Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

A Perspective on Data and Privacy for AI in Healthcare [Industrial and Governmental Activities]

As large language models continue to push the bounds of AI model size, they are also being trained on unprecedented volumes of data. While individual hospitals are estimated to produce petabytes of data per year, only a small fraction is currently being used for developing AI models. Additionally, with such data resources available, healthcare is well-positioned to benefit from the current trends in AI. Moreover, the inherently multi-modal and longitudinal nature of clinical data – from omics to imaging to unstructured notes – provides a fertile ground for the development and application of cutting-edge architectures like foundation models.

Gounley, John [Oak Ridge National Laboratory (ORNL↗

PeakQC: A Software Tool for Omics-Agnostic Automated Quality Control of Mass Spectrometry Data

Mass spectrometry is broadly employed to study complex molecular mechanisms in various biological and environmental fields, enabling 'omics' research such as proteomics, metabolomics, and lipidomics. As study cohorts grow larger and more complex with dozens to hundreds of samples, the need for robust quality control (QC) measures through automated software tools becomes paramount to ensure the integrity, high quality, and validity of scientific conclusions from downstream analyses and minimize the waste of resources. Since existing QC tools are mostly dedicated to proteomics, automated solutions supporting metabolomics are needed. To address this need, we developed the software PeakQC, a tool for automated QC of MS data that is independent of omics molecular types (i.e., omics-agnostic). It allows automated extraction and inspection of peak metrics of precursor ions (e.g., errors in mass, retention time, arrival time) and supports various instrumentations and acquisition types, from infusion experiments or using liquid chromatography and/or ion mobility spectrometry front-end separations and with/without fragmentation spectra from data-dependent or independent acquisition analyses. Diagnostic plots for fragmentation spectra are also generated. Here, in this paper, we describe and illustrate PeakQC’s functionalities using different representative data sets, demonstrating its utility as a valuable tool for enhancing the quality and reliability of omics mass spectrometry analyses.

47 OTHER INSTRUMENTATION↗

Model of metabolism and gene expression predicts proteome allocation in Pseudomonas putida

Abstract The genome-scale model of metabolism and gene expression (ME-model) forPseudomonas putidaKT2440,iPpu1676-ME, provides a comprehensive representation of biosynthetic costs and proteome allocation. Compared to a metabolic-only model,iPpu1676-ME significantly expands on gene expression, macromolecular assembly, and cofactor utilization, enabling accurate growth predictions without additional constraints. Multi-omics analysis using RNA sequencing and ribosomal profiling data revealed translational prioritization inP. putida, with core pathways, such as nicotinamide biosynthesis and queuosine metabolism, exhibiting higher translational efficiency, while secondary pathways displayed lower priority. Notably, the ME-model significantly outperformed the M-model in alignment with multi-omics data, thereby validating its predictive capacity. Thus,iPpu1676-ME offers valuable insights intoP. putida’s proteome allocation and presents a powerful tool for understanding resource allocation in this industrially relevant microorganism.

Mathematical & Computational Biology↗

RhizoGrid Indexed Sorghum Rhizosphere Multi-Omics

PerCon SFA project data dentification of spatially resolved biomarkers of drought in Sorghum bicolor rhizosphere molecular-microbe interactions using a novel root cartography "RhizoGrid" system for sampling plants under drought and control conditions across 10 equally sized root zone environments (4 quadrants each). Each quadrant was sampled and processed for 16S amplicon, metabolomics, and X-ray computed tomography (XCT). Data download includes experimental metadata and results files for 16S rRNA sequence analysis of microbial community assembly (processed data files), liquid chromatography mass spectrometry (LC-MS) metabolomics analysis of microbial community root exudates (processed data files), X-ray computed tomography (XCT) spatial gradient analysis (raw and processed data files) of microbial community composition, and related computational modeling outputs.

59 BASIC BIOLOGICAL SCIENCES↗

Bleach Rescues Nannochloropsis from an Obligate Parasite and Alters Microbial and Metabolite Signatures of Outdoor Cultures

Chemical agents are commonly used to protect algal crops. Yet, few studies have characterized the effects of these agents on associated microbial communities to understand effects on microbial functions relevant to algal crop production and protection. Here, we used shotgun metagenomic sequencing and untargeted exometabolite profiling to link the application of bleach, a -cidal agent used to protect algae from pests, to changes in community composition, metabolic pathways, and exometabolies - at a whole community level. Bleach protected the algal crop from crashing but altered bacterial diversity. Analysis of metagenome-assembled genomes (MAGs) revealed a classic predator-prey cycle between Oligoflexus and our target alga Nannochloropsis. Olifoflexus genomes from our study were notably similar to a previously identified BALO (Bdellovibrio and like organism), FD111, known to kill Nannochloropsis cultures, providing strong evidence that an FD111-like organism was responsible for the crash. Metabolic pathway composition differed between bleached and unbleached ponds, with abundance of twelve pathways related to stress tolerance, including the superpathway of methylglyoxal degradation, lipid IVA biosynthesis, and ectoine biosynthesis, greater in bleached ponds compared to unbleached ponds. Virulence factors related to adherence, biofilm formation, motility, and pathogenicity increased dramatically in bleached ponds with time, although this increase was not coupled with an increase in pathogens - algal or otherwise - or a decline in algal health. Our study highlights the importance of coupling 16S rRNA gene sequencing with whole genome data and other -omics tools to sketch a larger picture of community structure and function in crop systems. Moreover, our results highlight that continued long-term bleaching may lead to negative effects to crop health or downstream adverse health effects to humans or animals, depending on the algal product (i.e. human supplements or animal feedstocks). Future work on alternative treatment methods that would reduce resistance is necessary in the field.

09 BIOMASS FUELS↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

spammR: an R package designed for analysis and integration of spatial multi-omic measurements

Spatial omics is a young and evolving field and as such shows rapid development of novel technologies and analysis methods to measure transcripts, proteins, metabolites, and post-translational modifications at high spatial resolution. These advances in technology have enabled the simultaneous generation of abundance profiles for multiple different omics types and associated microscopy imaging data, as well as their analysis in a spatial context. However, most analytical tools are designed for spatial transcriptomics platforms and are challenging to use in other contexts such as mass spectrometry-based measurements or metagenomics. To this end we present spammR (spatial analysis of multi-omics measurements in R), an R package that enables end-to-end analysis with a specific focus on mass-spectrometry derived spatial omics datasets with (1) smaller sample sizes and spatial sparsity of samples, (2) considerable missingness, and (3) no a-priori knowledge about proteins or genes of interest, relying on a fully data-driven approach.

spammR↗

The Future of a Myriad of Accelerated Biodiscoveries Lies in AI‐Powered Mass Spectrometry and Multiomics Integration

The intersection of modern artificial intelligence (AI) and mass spectrometry (MS) is set to transform the MS‐based “omics” research fields, particularly proteomics, metabolomics, lipidomics, and glycomics, enabling advancements across a wide range of domains, from health to environment and industrial biotechnology. Beginning with an overview of key challenges inherent in MS software pipelines, this personal perspective explores how AI‐driven solutions can address them to enhance data processing, integration and interpretation. It proposes a paradigm shift in molecular identification and quantitation algorithms, leveraging AI to enable holistic interpretation of MS‐based multiomics data. While centered on MS‐based omics, this holistic AI‐driven paradigm is also critical for connecting dynamic biochemical changes to genomics and transcriptomics contexts, reinforcing the integrative value of MS in multiomics research. Ultimately, this AI‐driven approach could enhance efficiency, accuracy, and molecular breadth of coverage, deepening our systems‐level understanding of biological processes and accelerating a myriad of biodiscoveries.

47 OTHER INSTRUMENTATION↗

Metabolic rewiring and biomass redistribution enable optimized mixotrophic growth in Chlamydomonas

Aquatic photosynthetic systems account for approximately one-half of all global carbon assimilation and could be a significant source of renewable fuels and feedstocks. However, rapid growth and biomass production in algae have not always translated into high product yields, partly because central metabolism is context specific, with metabolic fluxes being influenced by nutrient conditions and other environmental factors. In the green microalga Chlamydomonas reinhardtii (Chlamydomonas), mixotrophic cultures (acetate + light) grow far faster than phototrophic (light only) or heterotrophic (acetate + dark) cultures, even though acetate partially suppresses photosynthesis. Here, an isotopic dilution strategy with unlabeled acetate was combined with 13 CO 2 transient labeling to perform isotopically nonstationary metabolic flux analysis (INST-MFA) and to directly compare autotrophic and mixotrophic metabolism in Chlamydomonas supported by data from transcriptomics, proteomics, and metabolomics. INST-MFA indicated that acetate induces a synergistic rewiring of metabolism, conserving carbon by using the glyoxylate cycle and suppressing gluconeogenesis, the latter of which was discordant with omics results and prior models. Additionally, our data provide a plausible rationale for the well-known suppression of photosynthesis by acetate. We propose that reduced total protein content in mixotrophic versus phototrophic cells, much of which is attributed to reduced levels of photosynthetic proteins, decreases the costly metabolic burden of protein synthesis and represents a growth rate optimization strategy.

59 BASIC BIOLOGICAL SCIENCES↗

Integrated multi-omic characterizations of the synapse reveal RNA processing factors and ubiquitin ligases associated with neurodevelopmental disorders

The molecular composition of the excitatory synapse is incompletely defined due to its dynamic nature across developmental stages and neuronal populations. To address this gap, we apply proteomic mass spectrometry to characterize the synapse in multiple biological models including the fetal human brain and hiPSC-derived neurons. To prioritize the identified proteins, we develop an orthogonal multi-omic screen of genomic, transcriptomic, interactomic, and structural data. This data-driven framework identifies proteins with key molecular features intrinsic to the synapse, including characteristic patterns of biophysical interactions and cross-tissue expression. The multi-omic analysis captures synaptic proteins across developmental stages and experimental systems, including 493 synaptic candidates supported by proteomics. We further investigate three such proteins that are associated with neurodevelopmental disorders – the CUL3 E3 ubiquitin ligase, the DDX3X and YBX1 nucleic-acid binding proteins – by mapping their networks of physically interacting synapse proteins or transcripts. Our study demonstrates the potential of an integrated multi-omic approach to systematically and more comprehensively resolve the synaptic architecture.

59 BASIC BIOLOGICAL SCIENCES↗

Single-cell and spatial omics in plants: from cellular atlases to regulatory mechanisms

Single-cell RNA sequencing (scRNA-seq) has transformed transcriptomic studies by enabling gene expression profiling at the resolution of individual cells within and across a broad range of tissue types, revealing cellular heterogeneity that is obscured in bulk tissue transcriptomes. Over the past decade, improvements in microfluidics and library preparation have drastically increased throughput, allowing tens of thousands of cells to be assayed in a single experiment. Although initially developed in animal systems, scRNA-seq has rapidly emerged as a powerful and widely adopted approach in plant biology. Beyond transcriptomics, the integration of single-cell data with chromatin accessibility, proteomics, metabolomics, and spatial omics is enabling a system-level understanding of plant gene regulation and cellular organization. Network-based analytical frameworks further support the reconstruction of gene regulatory networks and the interpretation of complex single-cell data. In this review, we summarize the current technological landscape of plant single-cell studies, discuss key experimental and analytical challenges, and review emerging strategies for validating single-cell discoveries. We also discuss future directions in applying single-cell technologies to woody perennials plants and bioenergy-relevant crops, emphasizing their potential to accelerate the discovery of cell type-specific regulatory mechanisms underlying growth, stress resilience, and biomass production.

Li, Miaomiao [ORNL] (ORCID:0000000321326168)↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Enhanced Spatial Proteomics and Metabolomics from a Single Tissue Section Using MALDI-MSI and LCM-microPOTS Platforms

Spatially resolved mass spectrometry (MS)-based multi-omics workflows are becoming more utilized for revealing the complex biology that occurs within tissues. However, these approaches commonly require multiple independent tissue sections to analyze the metabolite and protein compositions of these samples. This poses a significant challenge in preserving cell- or region-specific molecular fidelity, as variations between tissue sections can compromise the accurate correlation of molecular data. Here, in this study, we developed workflows for comprehensive multi-omics profiling from a single tissue section (STS) using different MS modalities. We enhanced the functionality of an electrically insulated substrate by employing metal-assisted approaches that enabled both MS-based untargeted spatial metabolomics and proteomics from STS. This allowed metabolite imaging using matrix-assisted laser desorption/ionization-MS imaging (MALDI-MSI), without compromising it for subsequent proteome profiling with laser capture microdissection (LCM)-based technology. Specifically, implementing copper tape as a backing for polyethylene naphthalate (PEN) slides enabled the detection of >140 metabolites across a poplar root tissue section using MALDI-trapped ion mobility spectrometry time of flight (timsTOF)-MS. Afterwards, we detected 6,571 unique proteins from two distinct root regions by leveraging LCM technology coupled to our microdroplet based sample preparation approach. We also developed an alternative workflow utilizing gold-coated PEN substrates for imaging with MALDI-Fourier-transform ion cyclotron resonance (FTICR)-MS, which permitted the profiling of >170 metabolites and the identification of 6,542 unique proteins across a single poplar root tissue section. These results were comparable to using each assay independently without modifications. These approaches offer new opportunities for high-resolution molecular profiling of multiple omics-levels across biological tissues.

Veličković, Marija [Pacific Northwest National Lab↗