Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Molecular formula”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13 C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13 C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13 C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13 C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.

Carbon↗

Computer programs for the interpretation of low resolution mass spectra: Program for calculation of molecular isotopic distribution and program for assignment of molecular formulas

Two FORTRAN computer programs for the interpretation of low resolution mass spectra were prepared and tested. One is for the calculation of the molecular isotopic distribution of any species from stored elemental distributions. The program requires only the input of the molecular formula and was designed for compatability with any computer system. The other program is for the determination of all possible combinations of atoms (and radicals) which may form an ion having a particular integer mass. It also uses a simplified input scheme and was designed for compatability with any system.

Miller, R. A.↗

Riverine organic matter functional diversity increases with catchment size

A large amount of dissolved organic matter (DOM) is transported to the ocean from terrestrial inputs each year (~0.95 Pg C per year) and undergoes a series of abiotic and biotic reactions, causing a significant release of CO 2 . Combined, these reactions result in variable DOM characteristics (e.g., nominal oxidation state of carbon, double-bond equivalents, chemodiversity) which have demonstrated impacts on biogeochemistry and ecosystem function. Despite this importance, however, comparatively few studies focus on the drivers for DOM chemodiversity along a riverine continuum. Here, we characterized DOM within samples collected from a stream network in the Yakima River Basin using ultrahigh-resolution mass spectrometry (i.e., FTICR-MS). To link DOM chemistry to potential function, we identified putative biochemical transformations within each sample. We also used various molecular characteristics (e.g., thermodynamic favorability, degradability) to calculate a series of functional diversity metrics. We observed that the diversity of biochemical transformations increased with increasing upstream catchment area and landcover. This increase was also connected to expanding functional diversity of the molecular formula. This pattern suggests that as molecular formulas become more diverse in thermodynamics or degradability, there is increased opportunity for biochemical transformations, potentially creating a self-reinforcing cycle where transformations in turn increase diversity and diversity increase transformations. We also observed that these patterns are, in part, connected to landcover whereby the occurrence of many landcover types (e.g., agriculture, urban, forest, shrub) could expand DOM functional diversity. For example, we observed that a novel functional diversity metric measuring similarity to common freshwater molecular formulas (i.e., carboxyl-rich alicyclic molecules) was significantly related to urban coverage. These results show that DOM diversity does not decrease along stream networks, as predicted by a common conceptual model known as the River Continuum Concept, but rather are influenced by the thermodynamic and degradation potential of molecular formula within the DOM, as well as landcover patterns.

54 ENVIRONMENTAL SCIENCES↗

Ultrahigh performance LC/FT-MS non-targeted screening for biomass burning organic aerosol with MZmine2 and MFAssignR

In recent years, ultrahigh performance liquid chromatography Fourier transform mass spectrometry (LC/FT-MS) based non-targeted screening (NTS) methods have become increasingly popular for comprehensive analysis of complex organic mixtures. However, applying these methods for environmental complex mixture analysis is challenging due to the extreme complexity of natural samples and a lack of standard samples or surrogates for environmental complex mixtures. Furthermore, limited molecular markers in the databases and insufficient data processing software workflows make the application of these methods more challenging for environmental complex mixtures. In this work, we implement a new NTS data processing workflow to process data collected from ultrahigh performance liquid chromatography and Fourier transform Orbitrap Elite Mass Spectrometry (LC/FT-MS) by combining MZmine2 and MFAssignR, two opensource data processing tools and commercial Mesquite liquid smoke as a surrogate for biomass burning organic aerosol. MZmine2.53 data extraction followed MFAssignR molecular formula assignment offered noise free and highly accurate 1733 individual molecular formulas presented in liquid smoke with 4906 molecular species, including isomers. The results of this new approach were consistent with the results of direct infusion FT-MS analysis confirming its reliability. Over 90% of the molecular formulas presented in mesquite liquid smoke were matched with the molecular formulas of ambient biomass burning organic aerosol. This suggests the potential use of commercial liquid smoke as a surrogate for biomass burning organic aerosol research. Furthermore, the presented method significantly improves the identification of the molecular composition of biomass burning organic aerosol by successfully addressing some of the limitations related to the data analysis and giving a semi quantitative insight into the analysis.

54 ENVIRONMENTAL SCIENCES↗

Soil microbiome resilience to short-term (30 days, 90 days) and long-term (1000 days) drought

This dataset contains data used for the paper "Drought duration does not impact soil microbiome resilience". The Related References will be updated with a full citation when available. Increasing global droughts exert large but poorly understood effects on the microbial communities and ecology of soil. Microbial communities generally show resilience and return to pre-drought conditions when short-term droughted soils are rewet; soils exposed to long-term drought, however, often show a lag upon rewetting, after which microbial communities may or may not return to their pre-stressed conditions. Though short-term droughts have been widely studied, long-term drought manipulation experiments remain rare, especially those that compare microbial response to short-term and long-term drought in tandem. We conducted a 1000-day drought simulation in controlled laboratory conditions with soil cores collected from a tidal freshwater ecosystem in Washington state, USA, and subsequently exposed them to rewetting for two weeks. We also included short-term (30-day and 90-day) drought and rewet treatments to directly compare microbial community and organic matter responses across drought durations. We found distinct microbial taxa belonging to Firmicutes and Actinobacteria enriched after the 1000-day drought, but not after the short-term droughts. While we hypothesized that the microbial community would recover from a short-term drought after rewetting to resemble pre-drought conditions, our results revealed community dissimilarities between rewet and pre-drought conditions across all drought durations. These findings suggest unique microbial life history strategies within certain microbial phyla that make them successful colonizers during an extended drought period, and the influence of environmental and physiological context on microbial responses to rewetting. The 16SrRNA gene amplicon dataset contains processed DNA sequences in the form of an ASV table with raw unrarefied read counts and representative sequences in .fasta format as described in the ESS-DIVE amplicon sequence reporting format (https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format/instructions). The Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) dataset consists of processed files containing presence absence data of molecular formulae and molecular characterization of FTICR resolved peaks. The Nuclear Magnetic Resonance (NMR) dataset contains files relevant to NMR spectra and peaks. A sample key file and a sample metadata file is included for the FTICR/NMR and 16S dataset respectively.

1000-day drought↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

IsoMatchMS : Open-Source Software for Automated Annotation and Visualization of High Resolution MALDI-MS Spectra

Due to its speed, accuracy, and adaptability to various sample types, matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) has become a popular method to identify molecular isotope profiles from biological samples. Often MALDI-MS data do not include tandem MS fragmentation data, and thus the identification of compounds in samples requires external databases so that the accurate mass of detected signals can be matched to known molecular compounds. Most relevant MALDI-MS software tools developed to confirm compound identifications are focused on small molecules (e.g., metabolites, lipids) and cannot be easily adapted to protein data due to their more complex isotopic distributions. Here, we present an R package called IsoMatchMS for the automated annotation of MALDI-MS data for multiple datatypes: intact proteins, peptides, and glycans. This tool accepts already derived molecular formulas or, for proteomics applications, can derive molecular formulas from a list of input peptides or proteins including proteins with post-translational modifications. In conclusion, visualization of all matched isotopic profiles is provided in a highly accessible HTML format called a trelliscope display, which allows users to filter and sort by several parameters such as match scores and the number of peaks matched. IsoMatchMS simplifies the annotation and visualization of MALDI-MS data for downstream analyses.

47 OTHER INSTRUMENTATION↗

Improved Characterization of Soil Organic Matter by Integrating FT-ICR MS, Liquid Chromatography Tandem Mass Spectrometry, and Molecular Networking: A Case Study of Root Litter Decay under Drought Conditions

Understanding of how soil organic matter (SOM) chemistry is altered in a changing climate has advanced considerably; however, most SOM components remain unidentified, impeding the ability to characterize a major fraction of organic matter and predict what types of molecules, and from which sources, will persist in soil. Here we present a novel approach to better characterize SOM extracts by integrating information from three types of analyses, and we deploy this method to characterize decaying root-detritus soil microcosms subjected to either drought or normal conditions. To observe broad differences in composition, we employed direct infusion Fourier-transform ion cyclotron resonance mass spectrometry (DI-FT-ICR MS). We complemented this with liquid chromatography tandem mass spectrometry (LC-MS/MS) to identify components by library matching. Since libraries contain only a small fraction of SOM components, we also used fragment spectral cosine similarity scores to relate unknowns and library matches through molecular networks. This integrated approach allowed us to corroborate DI-FT-ICR MS molecular formulas using library matches, which included fungal metabolites and related polyphenolic compounds. We also inferred structures of unknowns from molecular networks and improved LC-MS/MS annotation rates from ~5 to 35% by considering DI-FT-ICR MS molecular formula assignments. Under drought conditions, we found greater relative amounts of lignin-like vs condensed aromatic polyphenol formulas and lower average nominal oxidation state of carbon, suggesting reduced decomposition of SOM and/or microbes under stress. Our integrated approach provides a framework for enhanced annotation of SOM components that is more comprehensive than performing individual data analyses in parallel.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Linking Dissolved Organic Matter Composition to Landscape Properties in Wetlands Across the United States of America

Abstract Wetlands are integral to the global carbon cycle, serving as both a source and a sink for organic carbon. Their potential for carbon storage will likely change in the coming decades in response to higher temperatures and variable precipitation patterns. We characterized the dissolved organic carbon (DOC) and dissolved organic matter (DOM) composition from 12 different wetland sites across the USA spanning gradients in climate, landcover, sampling depth, and hydroperiod for comparison to DOM in other inland waters. Using absorption spectroscopy, parallel factor analysis modeling, and ultra‐high resolution mass spectroscopy, we identified differences in DOM sourcing and processing by geographic site. Wetland DOM composition was driven primarily by differences in landcover where forested sites contained greater aromatic and oxygenated DOM content compared to grassland/herbaceous sites which were more aliphatic and enriched in N and S molecular formulae. Furthermore, surface and porewater DOM was also influenced by properties such as soil type, organic matter content, and precipitation. Surface water DOM was relatively enriched in oxygenated higher molecular weight formulae representing HUP High O/C compounds than porewaters, whose DOM composition suggests abiotic sulfurization from dissolved inorganic sulfide. Finally, we identified a group of persistent molecular formulae (3,489) present across all sites and sampling depths (i.e., the signature of wetland DOM) that are likely important for riverine‐to‐coastal DOM transport. As anthropogenic disturbances continue to impact temperate wetlands, this study highlights drivers of DOM composition fundamental for understanding how wetland organic carbon will change, and thus its role in biogeochemical cycling.

Environmental Sciences & Ecology↗

Applying the core-satellite species concept: Characteristics of rare and common riverine dissolved organic matter

Introduction: Dissolved organic matter (DOM) composition varies over space and time, with a multitude of factors driving the presence or absence of each compound found in the complex DOM mixture. Compounds ubiquitously present across a wide range of river systems (hereafter termed core compounds) may differ in chemical composition and reactivity from compounds present in only a few settings (hereafter termed satellite compounds). Here, we investigated the spatial patterns in DOM molecular formulae presence (occupancy) in surface water and sediments across 97 river corridors at a continental scale using the “Worldwide Hydrobiogeochemical Observation Network for Dynamic River Systems—WHONDRS” research consortium. Methods: We used a novel data-driven approach to identify core and satellite compounds and compared their molecular properties identified with Fourier-transform ion cyclotron resonance mass spectrometry (FT-ICR MS). Results: In this work, we found that core compounds clustered around intermediate hydrogen/carbon and oxygen/carbon ratios across both sediment and surface water samples, whereas the satellite compounds varied widely in their elemental composition. Within surface water samples, core compounds were dominated by lignin-like formulae, whereas protein-like formulae dominated the core pool in sediment samples. In contrast, satellite molecular formulae were more evenly distributed between compound classes in both sediment and water molecules. Core compounds found in both sediment and water exhibited lower molecular mass, lower oxidation state, and a higher degree of aromaticity, and were inferred to be more persistent than global satellite compounds. Higher putative biochemical transformations were found in core than satellite compounds, suggesting that the core pool was more processed. Discussion: The observed differences in chemical properties of core and satellite compounds point to potential differences in their sources and contribution to DOM processing in river corridors. Overall, our work points to the potential of data-driven approaches separating rare and common compounds to reduce some of the complexity inherent in studying riverine DOM.

54 ENVIRONMENTAL SCIENCES↗

Data Set Analysis to Reduce Uncertainty in Formula Assignments of Ultrahigh Resolution Mass Spectra

Environmental samples contain a vast array of organic compounds with diverse elemental compositions and heteroatom content. Molecular formula assignments of ultrahigh resolution mass spectra (HRMS) hold promise for elucidating the molecular composition of these compounds. However, the need to account for an assortment of heteroatoms increases the uncertainty associated with individual assignments – and ultimately the ecological, biological, and biogeochemical insights gleaned from the assignments. To address this challenge, we introduce a formula assignment strategy that leverages HRMS data sets to improve assignment confidence, filter false assignments, and mitigate bias in assignment routines. The strategy, implemented using CoreMS, first identifies the highest confidence assignment for a recurring ion in a data set by assessing the mass accuracy and isotopologue similarity of all assignments to the ion across the data set. The second component of the strategy examines the consistency of mass errors for an assigned ion throughout a data set and flags formulas with statistically unlikely deviations in mass error. Here, we illustrate the application and utility of the strategy by comparing its results against documented misassignment patterns within a set of oceanographic samples that were measured with 21 T Fourier Transform Ion Cyclotron Resonance Mass Spectrometry. Because the efficacy of our strategy improves with data set size, it is particularly useful for enhancing assignment confidence in large HRMS data sets common in studies of environmental systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Relating Molecular Properties to the Persistence of Marine Dissolved Organic Matter with Liquid Chromatography–Ultrahigh-Resolution Mass Spectrometry

Marine dissolved organic matter (DOM) contains a complex mixture of small molecules that eludes rapid biological degradation. Spatial and temporal variations in the abundance of DOM reflect the existence of fractions that are removed from the ocean over different time scales, ranging from seconds to millennia. However, it remains unknown whether the intrinsic chemical properties of these organic components relate to their persistence. Here, we elucidate and compare the molecular compositions of distinct DOM fractions with different lability along a water column in the North Atlantic Gyre. Our analysis utilized ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry at 21 T coupled to liquid chromatography and a novel data pipeline developed in CoreMS that generates molecular formula assignments and metrics of isomeric complexity. Clustering analysis binned 14 857 distinct molecular components into groups that correspond to the depth distribution of semilabile, semirefractory, and refractory fractions of DOM. The more labile fractions were concentrated near the ocean surface and contained more aliphatic, hydrophobic, and reduced molecules than the refractory fraction, which occurred uniformly throughout the water column. These findings suggest that processes that selectively remove hydrophobic compounds, such as aggregation and particle sorption, contribute to variable removal rates of marine DOM.

54 ENVIRONMENTAL SCIENCES↗

Unsupervised learning of representative local atomic arrangements in molecular dynamics data

Molecular dynamics (MD) simulations present a data-mining challenge, given that they can generate a considerable amount of data but often rely on limited or biased human interpretation to examine their information content. By not asking the right questions of MD data we may miss critical information hidden within it. Here we combine dimensionality reduction (UMAP) and unsupervised hierarchical clustering (HDBSCAN) to quantitatively characterize prevalent coordination environments of chemical species within MD data. By focusing on local coordination, we significantly reduce the amount of data to be analyzed by extracting all distinct molecular formulas within a given coordination sphere. We then efficiently combine UMAP and HDBSCAN with alignment or shape-matching algorithms to partition these formulas into structural isomer families indicating their relative populations. The method was employed to reveal details of cation coordination in electrolytes based on molecular liquids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tethered balloon system and High-Resolution Mass Spectrometry Reveal Increased Organonitrates Aloft Compared to the Ground Level

Atmospheric particles play critical roles in climate. However, significant knowledge gaps remain regarding the vertically resolved organic molecular-level composition of atmospheric particles due to aloft sampling challenges. To address this, we use a tethered balloon system at the Southern Great Plains Observatory and high-resolution mass spectrometry to, respectively, collect and characterize organic molecular formulas (MF) in the ground level and aloft (up to 750 m) samples. We show that organic MF uniquely detected aloft were dominated by organonitrates (139 MF; 54% of all uniquely detected aloft MF). Organonitrates that were uniquely detected aloft featured elevated O/C ratios (0.73 ± 0.23) compared to aloft organonitrates that were commonly observed at the ground level (0.63 ± 0.22). Unique aloft organic molecular composition was positively associated with increased cloud coverage, increased aloft relative humidity (~40% increase compared to ground level), and decreased vertical wind variance. Furthermore, 29% of extremely low volatility organic compounds in the aloft sample were truly unique to the aloft sample compared to the ground level, emphasizing potential oligomer formation at higher altitudes. Overall, this study highlights the importance of considering vertically resolved organic molecular composition (particularly for organonitrates) and hypothesizes that aqueous phase transformations and vertical wind variance may be key variables affecting the molecular composition of aloft organic aerosol.

54 ENVIRONMENTAL SCIENCES↗