SEARCH · Engineering Papers
Results for “Metabolomics”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Longitudinal Time-series Analysis Of The Effects Of Long-duration Space Flight In Male And Female Astronauts Using A 1h-nmr-based Metabolomics Approach
Explore the source record for details and available documents.
Ionic Silver affects Lettuce Growth, Nutrient Uptake, Microbiome and Metabolome in a Hydroponics System
Explore the source record for details and available documents.
Assessing and Manipulating ‘Garnet Giant’ Brassica juncea Metabolomic Phenotypes through LED Lighting
Explore the source record for details and available documents.
Effects of Low Dose Space Radiation Exposures on the Splenic Metabolome
Explore the source record for details and available documents.
Metabolomics as a Truly Translational Tool for Precision Medicine
Explore the source record for details and available documents.
Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets
Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.
EvoNet: A phylogenomic and systems biology approach to identify genes underlying plant survival in marginal, low‐N soils
The DOE‐BER “EvoNet” project investigates the genetic and molecular basis of plant resilience in extreme environments. We do this by identifying key genes that enable “extreme survivor” species to thrive in the nitrogen-poor soils of Chile’s hyper-arid Atacama Desert. Our collections focus on 32 Atacama extremophile species, including seven grass species with potential biofuel applications. To identify genes-of-importance to survival we compared genomic and transcriptomic profiles of extremophile species that thrive in the Atacama to those of closely related “sister” species from nitrogen-rich arid and mesic regions of California. Deep RNA sequencing and de novo transcriptome assembly across these triplet species sets supported a phylogenomic framework for identifying positively selected genes associated with adaptive divergence. Our integrative analysis combined ecological and environmental data, metagenomics, evolutionary and systems biology, and metabolomics. This enabled us to create an unprecedented framework for systematically understanding how non-model plants have adapted to survive in extreme conditions. Our resulting database of positively selected ortholog groups in the extremophile plants offers promising targets for engineering crop and biofuel species with enhanced resilience to drought and extreme weather. Additionally, our newest dataset explores and exploits a complementary metabolomic approach. This new aspect provides innovative strategies to manipulate plant cell metabolism, further supporting efforts to improve agricultural productivity in the face of extreme climates. Importantly, our combined evolutionary- and metabolomic-based strategies focused on convergent patterns of adaptation, providing a genetic and metabolomic toolkit for improving crop and biofuel resilience across diverse plant species. Finally, our novel exploration of ecological and evolutionary dynamics delivered to the community a phylogenomic computational pipeline called “PhyloGeneious.” Our continued adaptations of this pipeline are publicly available to expedite evolutionary genomic research for future scientific discoveries. In total, our DOE-BER has provided genomic, metabolomic, and computational strategies to understand how extremophile plants provide evolutionary and physiological targets for improving agricultural and biofuel production.
Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI
A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.
Advanced multi-modal mass spectrometry imaging reveals functional differences of placental villous compartments at microscale resolution
The placenta is a complex and heterogeneous organ that links the mother and fetus, playing a crucial role in nourishing and protecting the fetus throughout pregnancy. Integrative spatial multi-omics approaches can provide a systems-level understanding of molecular changes underlying the mechanisms leading to the histological variations of the placenta during healthy pregnancy and pregnancy complications. Herein, we advance our metabolome-informed proteome imaging (MIPI) workflow to include lipidomic imaging, while also expanding the molecular coverage of metabolomic imaging by incorporating on-tissue chemical derivatization (OTCD). The improved MIPI workflow advances biomedical investigations by leveraging state-of-the-art molecular imaging technologies. Lipidome imaging identifies molecular differences between two morphologically distinct compartments of a placental villous functional unit, syncytiotrophoblast (STB) and villous core. Next, our advanced metabolome imaging maps villous functional units with enriched metabolomic activities related to steroid and lipid metabolism, outlining distinct molecular distributions across morphologically different villous compartments. Complementary proteome imaging on these villous functional units reveals a plethora of fatty acid- and steroid-related enzymes uniquely distributed in STB and villous core compartments. Integration across our advanced MIPI imaging modalities enables the reconstruction of active biological pathways of molecular synthesis and maternal-fetal signaling across morphologically distinct placental villous compartments with micrometer-scale resolution.
Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry
Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.
Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites
Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.
Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability
Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.
An ecological framework for microbial metabolites in the ocean ecosystem
The ocean microbe‐metabolite network involves thousands of individual metabolites that encompass a breadth of chemical diversity and biological functions. These microbial metabolites mediate biogeochemical cycles, facilitate ecological relationships, and impact ecosystem health. While analytical advancements have begun to illuminate such roles, a challenge in navigating the deluge of marine metabolomics information is to identify a subset of metabolites that have the greatest ecosystem impact. Here, we present an ecological framework to distill knowledge of fundamental metabolites that underpin marine ecosystems. We borrow terms from macroecology that describe important species, namely “dominant,” “keystone,” and “indicator” species, and apply these designations to metabolites within the ocean microbial metabolome. These selected metabolites may shape marine community structure, function, and health and provide focal points for enhanced study of microbe‐metabolite networks. Applying ecological concepts to marine metabolites provides a path to leverage metabolomics data to better describe and predict marine microbial ecosystems.
Omics-Based Comparison of Fungal Virulence Genes, Biosynthetic Gene Clusters, and Small Molecules in Penicillium expansum and Penicillium chrysogenum
Penicillium expansum is a ubiquitous pathogenic fungus that causes blue mold decay of apple fruit postharvest, and another member of the genus, Penicillium chrysogenum, is a well-studied saprophyte valued for antibiotic and small molecule production. While these two fungi have been investigated individually, a recent discovery revealed that P. chrysogenum can block P. expansum-mediated decay of apple fruit. To shed light on this observation, we conducted a comparative genomic, transcriptomic, and metabolomic study of two P. chrysogenum (404 and 413) and two P. expansum (Pe21 and R19) isolates. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for P. chrysogenum isolates, and different between P. expansum isolates. Further, the two P. chrysogenum genomes revealed secondary metabolite gene clusters that varied widely from P. expansum. This included the absence of an intact patulin gene cluster in P. chrysogenum, which corroborates the metabolomic data regarding its inability to produce patulin. Additionally, a core subset of P. expansum virulence gene homologues were identified in P. chrysogenum and were similarly transcriptionally regulated in vitro. Molecules with varying biological activities, and phytohormone-like compounds were detected for the first time in P. expansum while antibiotics like penicillin G and other biologically active molecules were discovered in P. chrysogenum culture supernatants. Our findings provide a solid omics-based foundation of small molecule production in these two fungal species with implications in postharvest context and expand the current knowledge of the Penicillium-derived chemical repertoire for broader fundamental and practical applications.
OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data
Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.
PeakQC: A Software Tool for Omics-Agnostic Automated Quality Control of Mass Spectrometry Data
Mass spectrometry is broadly employed to study complex molecular mechanisms in various biological and environmental fields, enabling 'omics' research such as proteomics, metabolomics, and lipidomics. As study cohorts grow larger and more complex with dozens to hundreds of samples, the need for robust quality control (QC) measures through automated software tools becomes paramount to ensure the integrity, high quality, and validity of scientific conclusions from downstream analyses and minimize the waste of resources. Since existing QC tools are mostly dedicated to proteomics, automated solutions supporting metabolomics are needed. To address this need, we developed the software PeakQC, a tool for automated QC of MS data that is independent of omics molecular types (i.e., omics-agnostic). It allows automated extraction and inspection of peak metrics of precursor ions (e.g., errors in mass, retention time, arrival time) and supports various instrumentations and acquisition types, from infusion experiments or using liquid chromatography and/or ion mobility spectrometry front-end separations and with/without fragmentation spectra from data-dependent or independent acquisition analyses. Diagnostic plots for fragmentation spectra are also generated. Here, in this paper, we describe and illustrate PeakQC’s functionalities using different representative data sets, demonstrating its utility as a valuable tool for enhancing the quality and reliability of omics mass spectrometry analyses.
Endophyte‐induced systemic spatial reprogramming of metabolism in Populus trichocarpa roots under drought
Beneficial, facultative endophytes help plants thrive in challenging environments by altering their host's metabolism, but how these cellular scale metabolic changes propagate to the systems biology scale is unknown. In this work, we employed a high-resolution chemical imaging approach to map metabolic changes at the Populus trichocarpa root-zone and cell-type levels combined with machine learning (ML) models to identify root metabolites and exudates that have predictive power over treatment class. We found that a nine-strain consortium of beneficial endophytes differentially altered the metabolome of droughted root tissues in a manner specific to cell type and root zone, with endophyte abundance showing a clear correlation to individual metabolites. Our study demonstrates that integrating spatial metabolomics with ML can reveal localized metabolic patterns linked to root–microbe interactions and generate novel hypotheses about underlying biological mechanisms.