Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cheminformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

35 records · Page 2

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

Internal calibration of transient kinetic data via machine learning

The temporal analysis of products (TAP) reactor provides a vast amount of transient kinetic information that may be used to describe a variety of chemical features including residence time distributions, kinetic coefficients, number of active sites, reaction mechanism, etc. However, as with any measurement device, the TAP reactor signal is convoluted with noise and drift is common. In order to reduce the uncertainty of the kinetic measurement and any derived parameters or mechanisms, proper preprocessing must be performed prior to any advanced type of analysis. This preprocessing includes baseline correction, i.e., a shift in the voltage response, and calibration, i.e., a scaling of the flux response based on prior experiments. The traditional methodology of preprocessing requires significant user discretion and reliance on separate calibration experiments that may drift over time. Herein we use machine learning techniques combined with physical constraints to understand the noise and drift that is being generated within and between experiments for enhancement of the chemical kinetic signal. As such, the proposed methodology demonstrates clear benefits over the traditional preprocessing approach by eliminating the need for separate calibration experiments or heuristic input from the user.

36 MATERIALS SCIENCE↗

Chemical Signatures of Seasonally Unique Anthropogenic Influences on Organic Aerosol Composition in the Central Amazon

Urbanization and fires perturb the quantities and composition of fine organic aerosol in the central Amazon, with ramifications for radiative forcing and public health. These disturbances include not only direct emissions of particulates and secondary organic aerosol (SOA) precursors, but also changes in the pathways through which biogenic precursors form SOA. The composition of ambient organic aerosol is complex and incompletely characterized, encompassing millions of potential structures relatively few of which have been synthesized and characterized. Through analysis of submicron aerosol samples from the Green Ocean Amazon (GoAmazon2014/5) field campaign by two-dimensional gas chromatography coupled with machine learning, ~1300 unique compounds were traced and characterized over two seasons. Fires and urban emissions produced chemically and interseasonally distinct impacts on product signatures, with only ~50% of compounds observed in both seasons. Seasonally unique populations point to the importance of aqueous processing in Amazonian aerosol ageing, but further mechanistic insights are impeded by limited product identity knowledge. Further, less than 10% of compounds were identifiable at an isomer-specific level. Overall, the findings (i) provide compositional characterization of anthropogenic influence on submicron organic aerosol in the Amazon, (ii) identify key season-to-season differences in chemical signatures, and (iii) highlight high-priority knowledge gaps in current speciated knowledge.

54 ENVIRONMENTAL SCIENCES↗

SIMILE enables alignment of tandem mass spectra with statistical significance

Abstract Interrelating small molecules according to their aligned fragmentation spectra is central to tandem mass spectrometry-based untargeted metabolomics. Current alignment algorithms do not provide statistical significance and compounds that have multiple delocalized structural differences and therefore often fail to have their fragment ions aligned. Here we align fragmentation spectra with both statistical significance and allowance for multiple chemical differences using Significant Interrelation of MS/MS Ions via Laplacian Embedding (SIMILE). SIMILE yields spectral alignment inferred structural connections in molecular networks that are not found with cosine-based scoring algorithms. In addition, it is now possible to rank spectral alignments based on p-values in the exploration of structural relationships between compounds and enhance the chemical connectivity that can be obtained with molecular networking.

59 BASIC BIOLOGICAL SCIENCES↗

Architector for high-throughput cross-periodic table 3D complex building

Abstract Rare-earth and actinide complexes are critical for a wealth of clean-energy applications. Three-dimensional (3D) structural generation and prediction for these organometallic systems remains a challenge, limiting opportunities for computational chemical discovery. Here, we introduce Architector , a high-throughput in-silico synthesis code for s-, p-, d-, and f-block mononuclear organometallic complexes capable of capturing nearly the full diversity of the known experimental chemical space. Beyond known chemical space, Architector performs in-silico design of new complexes including any chemically accessible metal-ligand combinations. Architector leverages metal-center symmetry, interatomic force fields, and tight binding methods to build many possible 3D conformers from minimal 2D inputs including metal oxidation and spin state. Over a set of more than 6,000 x-ray diffraction (XRD)-determined complexes spanning the periodic table, we demonstrate quantitative agreement between Architector-predicted and experimentally observed structures. Further, we demonstrate out-of-the box conformer generation and energetic rankings of non-minimum energy conformers produced from Architector , which are critical for exploring potential energy surfaces and training force fields. Overall, Architector represents a transformative step towards cross-periodic table computational design of metal complex chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Standardized multi-omics of Earth’s microbiomes reveals microbial and metabolite diversity

Despite advances in sequencing, lack of standardization makes comparisons across studies challenging and hampers insights into the structure and function of microbial communities across multiple habitats on a planetary scale. Here we present a multi-omics analysis of a diverse set of 880 microbial community samples collected for the Earth Microbiome Project. We include amplicon (16S, 18S, ITS) and shotgun metagenomic sequence data, and untargeted metabolomics data (liquid chromatography-tandem mass spectrometry and gas chromatography mass spectrometry). We used standardized protocols and analytical methods to characterize microbial communities, focusing on relationships and co-occurrences of microbially related metabolites and microbial taxa across environments, thus allowing us to explore diversity at extraordinary scale. In addition to a reference database for metagenomic and metabolomic data, we provide a framework for incorporating additional studies, enabling the expansion of existing knowledge in the form of an evolving community resource. We demonstrate the utility of this database by testing the hypothesis that every microbe and metabolite is everywhere but the environment selects. Our results show that metabolite diversity exhibits turnover and nestedness related to both microbial communities and the environment, whereas the relative abundances of microbially related metabolites vary and co-occur with specific microbial consortia in a habitat-specific manner. We additionally show the power of certain chemistry, in particular terpenoids, in distinguishing Earth’s environments (for example, terrestrial plant surfaces and soils, freshwater and marine animal stool), as well as that of certain microbes including Conexibacter woesei (terrestrial soils), Haloquadratum walsbyi (marine deposits) and Pantoea dispersa (terrestrial plant detritus). This Resource provides insight into the taxa and metabolites within microbial communities from diverse habitats across Earth, informing both microbial and chemical ecology, and provides a foundation and methods for multi-omics microbiome studies of hosts and the environment.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset of solution-based inorganic materials synthesis procedures extracted from the scientific literature

The development of a materials synthesis route is usually based on heuristics and experience. A possible new approach would be to apply data-driven approaches to learn the patterns of synthesis from past experience and use them to predict the syntheses of novel materials. However, this route is impeded by the lack of a large-scale database of synthesis formulations. In this work, we applied advanced machine learning and natural language processing techniques to construct a dataset of 35,675 solution-based synthesis procedures extracted from the scientific literature. Each procedure contains essential synthesis information including the precursors and target materials, their quantities, and the synthesis actions and corresponding attributes. Every procedure is also augmented with the reaction formula. Through this work, we are making freely available the first large dataset of solution-based inorganic materials synthesis procedures.

36 MATERIALS SCIENCE↗

Text-mined dataset of gold nanoparticle synthesis procedures, morphologies, and size entities

Abstract Gold nanoparticles are highly desired for a range of technological applications due to their tunable properties, which are dictated by the size and shape of the constituent particles. Many heuristic methods for controlling the morphological characteristics of gold nanoparticles are well known. However, the underlying mechanisms controlling their size and shape remain poorly understood, partly due to the immense range of possible combinations of synthesis parameters. Data-driven methods can offer insight to help guide understanding of these underlying mechanisms, so long as sufficient synthesis data are available. To facilitate data mining in this direction, we have constructed and made publicly available a dataset of codified gold nanoparticle synthesis protocols and outcomes extracted directly from the nanoparticle materials science literature using natural language processing and text-mining techniques. This dataset contains 5,154 data records, each representing a single gold nanoparticle synthesis article, filtered from a database of 4,973,165 publications. Each record contains codified synthesis protocols and extracted morphological information from a total of 7,608 experimental and 12,519 characterization paragraphs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Perovskite- and Dye-Sensitized Solar-Cell Device Databases Auto-generated Using ChemDataExtractor

The number of scientific publications reporting cutting-edge third-generation photovoltaic devices is increasing rapidly, owing to the pressing need to develop renewable-energy technologies that address the climate-change crisis. Consequently, the field could benefit from a central repository where photovoltaic-performance metrics, such as the power-conversion efficiency (η) are recorded. We present two automatically generated databases that contain photovoltaic properties and device material data for dye-sensitized solar cells (DSCs) and perovskite solar cells (PSCs), totalling 660,881 data entries representing 57,678 photovoltaic devices. The databases were generated by applying the text-mining toolkit ChemDataExtractor on a corpus of 25,720 articles. A multi-faceted evaluation, incorporating manual and automatic methods, was applied to ensure that the data contained therein were of the highest quality, with precision metrics ranging from 73.1% to 95.8%. The DSC database contains 475,045 entries representing 41,680 devices, and the PSC database contains 185,836 entries representing 15,818 devices. The databases are available in MongoDB and JSON formats, which can be queried in Python, R, Java and MATLAB for data-driven photovoltaic materials discovery.

14 SOLAR ENERGY↗

High accuracy barrier heights, enthalpies, and rate coefficients for chemical reactions

Abstract Quantitative chemical reaction data, including activation energies and reaction rates, are crucial for developing detailed kinetic mechanisms and accurately predicting reaction outcomes. However, such data are often difficult to find, and high-quality datasets are especially rare. Here, we use CCSD(T)-F12a/cc-pVDZ-F12// ω B97X-D3/def2-TZVP to obtain high-quality single point calculations for nearly 22,000 unique stable species and transition states. We report the results from these quantum chemistry calculations and extract the barrier heights and reaction enthalpies to create a kinetics dataset of nearly 12,000 gas-phase reactions. These reactions involve H, C, N, and O, contain up to seven heavy atoms, and have cleaned atom-mapped SMILES. Our higher-accuracy coupled-cluster barrier heights differ significantly (RMSE of ∼5 kcal mol −1 ) relative to those calculated at ω B97X-D3/def2-TZVP. We also report accurate transition state theory rate coefficients $${k}_{\infty }(T)$$ k ∞ ( T ) between 300 K and 2000 K and the corresponding Arrhenius parameters for a subset of rigid reactions. We believe this data will accelerate development of automated and reliable methods for quantitative reaction prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A thermoelectric materials database auto-generated from the scientific literature using ChemDataExtractor

An auto-generated thermoelectric-materials database is presented, containing 22,805 data records, automatically generated from the scientific literature, spanning 10,641 unique extracted chemical names. Each record contains a chemical entity and one of the seminal thermoelectric properties: thermoelectric figure of merit, ZT; thermal conductivity, κ; Seebeck coefficient, S; electrical conductivity, σ; power factor, PF; each linked to their corresponding recorded temperature, T. The database was auto-generated using the automatic sentence-parsing capabilities of the chemistry-aware, natural language processing toolkit, ChemDataExtractor 2.0, adapted for application in the thermoelectric-materials domain, following a rule-based sentence-simplification step. Data were mined from the text of 60,843 scientific papers that were sourced from three scientific publishers: Elsevier, the Royal Society of Chemistry, and Springer. To the best of our knowledge, this is the first automatically-generated database of thermoelectric materials and their properties from existing literature. The database was evaluated to have a precision of 82.25% and has been made publicly available to facilitate the application of data science in the thermoelectric-materials domain, for analysis, design, and prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A database of thermally activated delayed fluorescent molecules auto-generated from scientific literature with ChemDataExtractor

A database of thermally activated delayed fluorescent (TADF) molecules was automatically generated from the scientific literature. It consists of 25,482 data records with an overall precision of 82%. Among these, 5,349 records have chemical names in the form of SMILES strings which are represented with 91% accuracy; these are grouped in a subsidiary database. Each data record contains one of the following four properties: maximum emission wavelength (λ EM ), photoluminescence quantum yield (PLQY), singlet-triplet energy splitting (ΔE ST ), and delayed lifetime (τ D ). The databases were created through text mining using ChemDataExtractor, a chemistry-aware natural-language-processing toolkit, which has been adapted for TADF research. The text-mined corpus consisted of 2,733 papers from the Royal Society of Chemistry and Elsevier. To the best of our knowledge, these databases are the first databases that have been auto-generated for TADF molecules from existing publications. The databases have been publicly released for experimental and computational applications in the TADF research field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving deep learning model performance under parametric constraints for materials informatics applications

Abstract Modern machine learning (ML) and deep learning (DL) techniques using high-dimensional data representations have helped accelerate the materials discovery process by efficiently detecting hidden patterns in existing datasets and linking input representations to output properties for a better understanding of the scientific phenomenon. While a deep neural network comprised of fully connected layers has been widely used for materials property prediction, simply creating a deeper model with a large number of layers often faces with vanishing gradient problem, causing a degradation in the performance, thereby limiting usage. In this paper, we study and propose architectural principles to address the question of improving the performance of model training and inference under fixed parametric constraints. Here, we present a general deep-learning framework based on branched residual learning (BRNet) with fully connected layers that can work with any numerical vector-based representation as input to build accurate models to predict materials properties. We perform model training for materials properties using numerical vectors representing different composition-based attributes of the respective materials and compare the performance of the proposed models against traditional ML and existing DL architectures. We find that the proposed models are significantly more accurate than the ML/DL models for all data sizes by using different composition-based attributes as input. Further, branched learning requires fewer parameters and results in faster model training due to better convergence during the training phase than existing neural networks, thereby efficiently building accurate models for predicting materials properties.

36 MATERIALS SCIENCE↗

Deep learning workflow for the inverse design of molecules with specific optoelectronic properties

The inverse design of novel molecules with a desirable optoelectronic property requires consideration of the vast chemical spaces associated with varying chemical composition and molecular size. First principles-based property predictions have become increasingly helpful for assisting the selection of promising candidate chemical species for subsequent experimental validation. However, a brute-force computational screening of the entire chemical space is decidedly impossible. To alleviate the computational burden and accelerate rational molecular design, we here present an iterative deep learning workflow that combines (i) the density-functional tight-binding method for dynamic generation of property training data, (ii) a graph convolutional neural network surrogate model for rapid and reliable predictions of chemical and physical properties, and (iii) a masked language model. As proof of principle, we employ our workflow in the iterative generation of novel molecules with a target energy gap between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO).

97 MATHEMATICS AND COMPUTING↗

Autonomous design of new chemical reactions using a variational autoencoder

Artificial intelligence based chemistry models are a promising method of exploring chemical reaction design spaces. However, training datasets based on experimental synthesis are typically reported only for the optimal synthesis reactions. This leads to an inherited bias in the model predictions. Therefore, robust datasets that span the entirety of the solution space are necessary to remove inherited bias and permit complete training of the space. In this study, an artificial intelligence model based on a Variational AutoEncoder (VAE) has been developed and investigated to synthetically generate continuous datasets. The approach involves sampling the latent space to generate new chemical reactions. This developed technique is demonstrated by generating over 7,000,000 new reactions from a training dataset containing only 7,000 reactions. The generated reactions include molecular species that are larger and more diverse than the training set.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bioplastic design using multitask deep neural networks

Abstract Non-degradable plastic waste jeopardizes our environment, yet our modern lifestyle and current technologies are impossible to sustain without plastics. Bio-synthesized and biodegradable alternatives such as polyhydroxyalkanoates (PHAs) have the potential to replace large portions of the world’s plastic supply with cradle-to-cradle materials, but their chemical complexity and diversity limit traditional resource-intensive experimentation. Here, we develop multitask deep neural network property predictors using available experimental data for a diverse set of nearly 23,000 homo- and copolymer chemistries. Using the predictors, we identify 14 PHA-based bioplastics from a search space of almost 1.4 million candidates which could serve as potential replacements for seven petroleum-based commodity plastics that account for 75% of the world’s yearly plastic production. We also discuss possible synthesis routes for the identified promising materials.

36 MATERIALS SCIENCE↗