Engineering Papers⌕ Search

Engineering topics

Bilbao, Aivett

Publications and source records attributed to Bilbao, Aivett.

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhanced Detection of Primary Biological Aerosol Particles Using Machine Learning and Single-Particle Measurement

Accurately identifying primary biological aerosol particles (PBAPs) using analytical techniques poses inherent challenges due to their resemblance to other atmospheric carbonaceous particles. Here, we present a study of an enhanced method for detecting PBAPs by combining single-particle measurement with advanced supervised machine learning (SML) techniques. We analyzed ambient particles from a variety of environments and lab-generated standards, focusing on chemical composition for traditional rule-based and clustering approaches and incorporating morphological features into the SML approaches, neural networks and XGBoost, for improved accuracy. This study demonstrates that SML methods outperform traditional methods in quantifying PBAPs, achieving significant improvements in precision, recall, F1-score, and accuracy, leading to an increased number of detected PBAPs by at least 19%. The adaptability of the proposed XGBoost-based SML model is showcased in comparison to traditional methods in categorizing PBAPs for blind data sets from different geographical locations. Two field case studies were investigated, over agricultural land and Amazonia rain forest, representing relatively low and high concentrations of PBAPs, respectively, where XGBoost consistently detected up to 3.5 times more PBAPs than traditional methods. Precise detection of PBAPs in the atmosphere could significantly improve the prediction of climatic impacts by them.

42 ENGINEERING↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PeakQC: A Software Tool for Omics-Agnostic Automated Quality Control of Mass Spectrometry Data

Mass spectrometry is broadly employed to study complex molecular mechanisms in various biological and environmental fields, enabling 'omics' research such as proteomics, metabolomics, and lipidomics. As study cohorts grow larger and more complex with dozens to hundreds of samples, the need for robust quality control (QC) measures through automated software tools becomes paramount to ensure the integrity, high quality, and validity of scientific conclusions from downstream analyses and minimize the waste of resources. Since existing QC tools are mostly dedicated to proteomics, automated solutions supporting metabolomics are needed. To address this need, we developed the software PeakQC, a tool for automated QC of MS data that is independent of omics molecular types (i.e., omics-agnostic). It allows automated extraction and inspection of peak metrics of precursor ions (e.g., errors in mass, retention time, arrival time) and supports various instrumentations and acquisition types, from infusion experiments or using liquid chromatography and/or ion mobility spectrometry front-end separations and with/without fragmentation spectra from data-dependent or independent acquisition analyses. Diagnostic plots for fragmentation spectra are also generated. Here, in this paper, we describe and illustrate PeakQC’s functionalities using different representative data sets, demonstrating its utility as a valuable tool for enhancing the quality and reliability of omics mass spectrometry analyses.

47 OTHER INSTRUMENTATION↗

Evaluation of a Reference-Free Collision Cross Section Calibration Strategy for Proteomics Using SLIM-Based High-Resolution Ion Mobility Spectrometry–Mass Spectrometry

Ion mobility spectrometry (IMS) is a gas-phase analytical technique that separates ions with different sizes and shapes and is compatible with mass spectrometry (MS) to provide an additional separation dimension. The rapid nature of the IMS separation combined with the high sensitivity of MS-based detection and the ability to derive structural information on analytes in the form of the property collision cross section (CCS) makes IMS particularly well-suited for characterizing complex samples in -omics applications. In such applications, the quality of CCS from IMS measurements is critical to confident annotation of the detected components in the complex -omics samples. However, most IMS instrumentation in mainstream use requires calibration to calculate CCS from measured arrival times, with the most notable exception being drift tube IMS measurements using multifield methods. The strategy for calibrating CCS values, particularly selection of appropriate calibrants, has important implications for CCS accuracy, reproducibility, and transferability between laboratories. The conventional approach to CCS calibration involves explicitly defining calibrants ahead of data acquisition and crucially relies upon availability of reference CCS values. In this work, we present a novel reference-free approach to CCS calibration which leverages trends among putatively identified features and computational CCS prediction to conduct calibrations post-data acquisition and without relying on explicitly defined calibrants. We demonstrated the utility of this reference-free CCS calibration strategy for proteomics application using high-resolution structures for lossless ion manipulations (SLIM)-based IMS-MS. In conclusion, we first validated the accuracy of CCS values using a set of synthetic peptides and then demonstrated using a complex peptide sample from cell lysate.

59 BASIC BIOLOGICAL SCIENCES↗

Computational tools and algorithms for ion mobility spectrometry-mass spectrometry

Ion mobility spectrometry-mass spectrometry (IMS-MS or IM-MS) is a powerful analytical technique that combines the gas-phase separation capabilities of IM with the identification and quantification capabilities of MS. IM-MS can differentiate molecules with indistinguishable masses but different structures (e.g., isomers, isobars, molecular classes, and contaminant ions). The importance of this analytical technique is reflected by a staged increase in the number of applications for molecular characterization across a variety of fields, from different MS-based omics (proteomics, metabolomics, lipidomics, etc.) to the structural characterization of glycans, organic matter, proteins, and macromolecular complexes. With the increasing application of IM-MS there is a pressing need for effective and accessible computational tools. This article presents an overview of the most recent free and open-source software tools specifically tailored for the analysis and interpretation of data derived from IM-MS instrumentation. This review enumerates these tools and outlines their main algorithmic approaches, while highlighting representative applications across different fields. Finally, a discussion of current limitations and expectable improvements is presented.

59 BASIC BIOLOGICAL SCIENCES↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping microhabitats of lignocellulose decomposition by a microbial consortium

The leaf-cutter ant fungal garden ecosystem is a naturally evolved model system for efficient plant biomass degradation. Degradation processes mediated by the symbiotic fungus Leucoagaricus gongylophorus are difficult to characterize due to dynamic metabolisms and spatial complexity of the system. Herein, we performed microscale imaging across 12-µm-thick adjacent sections of Atta cephalotes fungal gardens and applied a metabolome-informed proteome imaging approach to map lignin degradation. This approach combines two spatial multiomics mass spectrometry modalities that enabled us to visualize colocalized metabolites and proteins across and through the fungal garden. Spatially profiled metabolites revealed an accumulation of lignin-related products, outlining morphologically unique lignin microhabitats. Metaproteomic analyses of these microhabitats revealed carbohydrate-degrading enzymes, indicating a prominent fungal role in lignocellulose decomposition. Integration of metabolome-informed proteome imaging data provides a comprehensive view of underlying biological pathways to inform our understanding of metabolic fungal pathways in plant matter degradation within the micrometer-scale environment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms and file structures to enhance software workflows for ion mobility mass spectrometry (IM-MS)

Support customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based metabolite characterization. Evaluate and improve the integration of ion mobility to existing MS analysis methods of the Mass Profiler Professional workflow (Mass Profiler, ID Browser and Mass Profiler Professional).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Metabolomic, photoprotective, and photosynthetic acclimatory responses to post‐flowering drought in sorghum

Abstract Climate change is globally affecting rainfall patterns, necessitating the improvement of drought tolerance in crops. Sorghum bicolor is a relatively drought‐tolerant cereal. Functional stay‐green sorghum genotypes can maintain green leaf area and efficient grain filling during terminal post‐flowering water deprivation, a period of ~10 weeks. To obtain molecular insights into these characteristics, two drought‐tolerant genotypes, BTx642 and RTx430, were grown in replicated control and terminal post‐flowering drought field plots in California's Central Valley. Photosynthetic, photoprotective, and water dynamics traits were quantified and correlated with metabolomic data collected from leaves, stems, and roots at multiple timepoints during control and drought conditions. Physiological and metabolomic data were then compared to longitudinal RNA sequencing data collected from these two genotypes. The unique metabolic and transcriptomic response to post‐flowering drought in sorghum supports a role for the metabolite galactinol in controlling photosynthetic activity through regulating stomatal closure in post‐flowering drought. Additionally, in the functional stay‐green genotype BTx642, photoprotective responses were specifically induced in post‐flowering drought, supporting a role for photoprotection in the molecular response associated with the functional stay‐green trait. From these insights, new pathways are identified that can be targeted to maximize yields under growth conditions with limited water.

54 ENVIRONMENTAL SCIENCES↗

Agilent CRADA (Abstract)

The CRADA between Agilent Technologies Inc. and Battelle will focus on five software components as listed below: Prototype 4D Feature Finding functionality with a particular focus on recovering low level features and extending the bottom end dynamic range of IM-MS technology. Compare and contrast developments to current 4D Feature Finding capabilities. Highlight important algorithmic aspects employed. Implement the PNNL saturation correction algorithm. Agilent will give PNNL the needed data file access API and assistance in understanding it implementation and any needed instrumental aspects. Supported high resolution products to include Agilent’s TOF, QTOF and IM-QTOF mass spectrometers. PNNL will then work with Agilent to benchmark performance. Implementation of the PNNL Hadamard de-multiplexing algorithm. Agilent will give provide PNNL the needed date file access API access and as needed assistance in understanding the current Agilent multiplexed IM offering. PNNL will then work with Agilent on benchmark performance. Add ion mobility collision cross sections to existing and new metabolomic libraries for data analysis with Agilent’s informatics program MPP/ID Browser. PNNL will work with Agilent to create a software pipeline that takes data from chemical and metabolic standards and properly formats it for inclusion in MPP accessible libraries, using the collision cross section as a new separation dimension. Improvements of MPP multidimensional matching to identify metabolomic features using multiple characteristics beyond retention time and accurate mass. Most significantly matching will include analyte collision cross section with proposed support for sample fraction or RapidFire cartridge and fragmentation spectra. PNNL will work with Agilent to modify and improve the current MPP analysis pipeline to allow for creating, aligning, and identifying MS features defined by accurate mass, collision cross section and chromatographic retention time. As additional criteria such as fraction or RapidFire cartridge type are supported in the identification process, then they also will become part of the automation workflow. This includes the automation of said system to work with command line program (i.e. not a GUI) sufficient for programmatic execution in a pipeline.

97 MATHEMATICS AND COMPUTING↗

Molecular Hypernetworks for Exploration of Multi-Dimensional Metabolomics Data (Chyper)

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and help address the challenge of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used, for example, in the Global Natural Products Social Molecular Networking platform, are an increasingly popular computational strategy for exploring and visualizing molecular relationships and improving annotation. MNs use graph representations to show the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated to a single molecular identity. However, more advanced methods may better represent the complexity present in samples. Our work aims to increase confidence in annotation propagation by extending molecular network methods to include “molecular hypernetworks” (MHNs), able to natively represent multiway relationships among observations supporting both human and analytical processing. In this paper we first introduce MHNs illustrated with simple examples, and demonstrate how to build them from liquid chromatography- and ion mobility spectrometry- separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

59 BASIC BIOLOGICAL SCIENCES↗