Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “molecular descriptor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Converting tabular data into images for deep learning with convolutional neural networks

Abstract Convolutional neural networks (CNNs) have been successfully used in many applications where important information about data is embedded in the order of features, such as speech and imaging. However, most tabular data do not assume a spatial relationship between features, and thus are unsuitable for modeling using CNNs. To meet this challenge, we develop a novel algorithm, image generator for tabular data (IGTD), to transform tabular data into images by assigning features to pixel positions so that similar features are close to each other in the image. The algorithm searches for an optimized assignment by minimizing the difference between the ranking of distances between features and the ranking of distances between their assigned pixels in the image. We apply IGTD to transform gene expression profiles of cancer cell lines (CCLs) and molecular descriptors of drugs into their respective image representations. Compared with existing transformation methods, IGTD generates compact image representations with better preservation of feature neighborhood structure. Evaluated on benchmark drug screening datasets, CNNs trained on IGTD image representations of CCLs and drugs exhibit a better performance of predicting anti-cancer drug response than both CNNs trained on alternative image representations and prediction models trained on the original tabular data.

59 BASIC BIOLOGICAL SCIENCES↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Artificial Neural Network Models for Octane Number and Octane Sensitivity: A Quantitative Structure Property Relationship Approach to Fuel Design

Octane sensitivity (OS), defined as the research octane number (RON) minus the motor octane number (MON) of a fuel, has gained interest among researchers due to its effect on knocking conditions in internal combustion engines. Compounds with a high OS enable higher efficiencies, especially within advanced compression ignition engines. RON/MON must be experimentally tested to determine OS, requiring time, funding, and specialized equipment. Thus, predictive models trained with existing experimental data and molecular descriptors (via quantitative structure-property relationships (QSPRs)) would allow for the preemptive screening of compounds prior to performing these experiments. Here, the present work proposes two methods for predicting the OS of a given compound: using artificial neural networks (ANNs) trained with QSPR descriptors to predict RON and MON individually to compute OS (derived octane sensitivity (dOS)), and using ANNs trained with QSPR descriptors to directly predict OS. Twenty-five ANNs were trained for both RON and MON and their test sets achieved an overall 6.4% and 5.2% error, respectively. Twenty-five additional ANNs were trained for both dOS and OS; dOS calculations were found to have 15.3% error while predicting OS directly resulted in 9.9% error. A chemical analysis of the top QSPR descriptors for RON/MON and OS is conducted, highlighting desirable structural features for high-performing molecules and offering insight into the inner mathematical workings of ANNs; such chemical interpretations study the interconnections between structural features, descriptors, and fuel performance showing that connectivity, structural diversity, and atomic hybridization consistently drive fuel performance.

09 BIOMASS FUELS↗

Developing a SARS-CoV-2 main protease binding prediction random forest model for drug repurposing for COVID-19 treatment

The coronavirus disease 2019 (COVID-19) global pandemic resulted in millions of people becoming infected with the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus and close to seven million deaths worldwide. It is essential to further explore and design effective COVID-19 treatment drugs that target the main protease of SARS-CoV-2, a major target for COVID-19 drugs. In this study, machine learning was applied for predicting the SARS-CoV-2 main protease binding of Food and Drug Administration (FDA)-approved drugs to assist in the identification of potential repurposing candidates for COVID-19 treatment. Ligands bound to the SARS-CoV-2 main protease in the Protein Data Bank and compounds experimentally tested in SARS-CoV-2 main protease binding assays in the literature were curated. These chemicals were divided into training (516 chemicals) and testing (360 chemicals) data sets. To identify SARS-CoV-2 main protease binders as potential candidates for repurposing to treat COVID-19, 1188 FDA-approved drugs from the Liver Toxicity Knowledge Base were obtained. A random forest algorithm was used for constructing predictive models based on molecular descriptors calculated using Mold2 software. Model performance was evaluated using 100 iterations of fivefold cross-validations which resulted in 78.8% balanced accuracy. The random forest model that was constructed from the whole training dataset was used to predict SARS-CoV-2 main protease binding on the testing set and the FDA-approved drugs. Model applicability domain and prediction confidence on drugs predicted as the main protease binders discovered 10 FDA-approved drugs as potential candidates for repurposing to treat COVID-19. Our results demonstrate that machine learning is an efficient method for drug repurposing and, thus, may accelerate drug development targeting SARS-CoV-2.

Research & Experimental Medicine↗

Advanced Modeling and Process-Materials Co-Optimization Strategies for Swing Adsorption Based Gas Separations

This project devised a computational framework for simultaneously co-optimizing pressure swing adsorption process designs along with the sorbent materials (specifically, metal-organic frameworks) to be employed in the associated packed bed columns. The materials optimization aspect involved search over a design space that can describe the material’s molecular structure, while the process optimization aspect considered various process degrees of freedom for steps arising in various cycle configurations. This framework was demonstrated on the separation of nitrogen and carbon dioxide, which arises ubiquitously in a multitude of post-combustion carbon capture and “blue” hydrogen production applications. Our results led to metal-organic framework molecular descriptor choices that are predicted to outperform standard structures used in practice, providing guidance for future metal-organic framework synthesis efforts.

20 FOSSIL-FUELED POWER PLANTS↗

From Structured Solvents to Hybrid Materials (SS2HM) for Chemically Selective Capture and Electromagnetic Release of CO 2 : Mechanisms, Stability and Interfaces (Final Report)

The goal of this research program was to develop high capacity sorbents amenable for alternative regeneration approaches for direct air capture (DAC) of CO 2 . In particular, the research aimed to develop an understanding of CO 2 binding mechanism, thermal and oxidative stability, and regeneration energetics of functionalized ionic liquids (ILs), deep eutectic solvents (DESs), and porous materials. ILs and DESs are high-dielectric solvents with structural tunability that permits the rational-design for energy-efficient regeneration approaches based on electromagnetic (EM) field and moisture-swing. By further incorporating these solvents into polymeric capsules and other structural supports, multi-scale interfaces for targeted CO 2 and energy transfers were achieved. Aspects related to CO 2 capacity, selectivity, stability, dielectric properties, and binding energies were examined through experimental and computational design to identify molecular descriptors to inform future design of structured solvents and hybrid materials for DAC. Enclosed final report details the key findings, science advancements, and workforce development efforts from this project.

36 MATERIALS SCIENCE↗

Finch: Toxicity Dose Response Curve Prediction of Chemical Compounds and Mixtures

A paradigm shift in chemical risk assessment is emphasizing mixture testing over single compound analysis, eliminating animal testing, and adopting advanced modeling approaches to understand mixture activity profiles. However, existing computational models largely focus on single chemicals, with few effective solutions for modeling complex mixtures that account for synergistic or antagonistic effects and multiple Modes of Action (MoA). Conventional methods like concentration addition (CA) and independent action (IA) are insufficient for this task as they are designed for simplistic interactions and struggle to account for the dynamic and multifaceted nature of chemical mixtures, such as overlapping MoA and non-linear interactions. Finch offers a novel approach utilizing deep learning (DL) embeddings and multi-task quantitative structure-activity relationship (QSAR) models to improve chemical exposure prediction. By leveraging molecular descriptors, physiochemical properties, and large language model (LLM) embeddings from SMILES inputs, Finch preserves critical information in a latent space thereby enhancing predictive accuracy. The multi-task learning aspect of Finch is highly advantageous, as it simultaneously optimizes multiple loss functions, leveraging all available data across tasks to develop generalized representations that effectively capture complex ingredient interactions within mixtures.

59 BASIC BIOLOGICAL SCIENCES↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

E min : A First-Principles Thermochemical Descriptor for Predicting Molecular Synthesizability

Predicting the synthesizability of a new molecule remains an unsolved challenge that chemists have long tackled with heuristic approaches. Here, in this study, we report a new method for predicting synthesizability using a simple yet accurate thermochemical descriptor. We introduce E min , the energy difference between a molecule and its lowest energy constitutional isomer, as a synthesizability predictor that is accurate, physically meaningful, and first-principles based. We apply E min to 134,000 molecules in the QM9 data set and find that E min is accurate when used alone and reduces incorrect predictions of "synthesizable" by up to 52% when used to augment commonly used prediction methods. Our work illustrates how first-principles thermochemistry and heuristic approximations for molecular stability are complementary, opening a new direction for synthesizability prediction methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning for Prediction of Thermodynamic Descriptors

Our objective is to apply machine learning (ML) algorithms for the prediction of molecular catalysis descriptors from geometric properties derived from experimental crystallographic databases. Catalysis is often considered a “low-data” discipline that is poorly suited for ML methods. An exception is the extensive structural information that is available for molecular catalysts through the Cambridge Structural Database (CSD), which contains atomically precise molecular structures from X-ray diffraction analysis for >600K metal complexes. As a proof-of-principle, we targeted the prediction of hydricity, a thermodynamic property that provides understanding and control of catalytic hydride transfer. We built a training set composed of ~100 molecular complexes with a known hydricity and structural information from the CSD. This data set was converted into a machine-readable format using the smooth overlap of atomic positions (SOAP) representation and further labeled with simple electronic descriptors for the metal centers. Multiple different neural networks were trained on this data set, and the accuracy of the hydricity predictions ranged from < 2 kcal/mol to 20 kcal/mol. The accuracy of each model was highly sensitive to which compounds were in the train versus test set, underscoring the challenges associated with small and chemically diverse data sets. Finally, to further augment the data set, we attempted to experimentally measure several new hydricity values, however these experiments were unsuccessful due to undesired chemical reactivity of the selected complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integration of Computational Docking into Anti-Cancer Drug Response Prediction Models

Cancer is a heterogeneous disease in that tumors of the same histology type can respond differently to a treatment. Anti-cancer drug response prediction is of paramount importance for both drug development and patient treatment design. Although various computational methods and data have been used to develop drug response prediction models, it remains a challenging problem due to the complexities of cancer mechanisms and cancer-drug interactions. To better characterize the interaction between cancer and drugs, we investigate the feasibility of integrating computationally derived features of molecular mechanisms of action into prediction models. Specifically, we add docking scores of drug molecules and target proteins in combination with cancer gene expressions and molecular drug descriptors for building response models. The results demonstrate a marginal improvement in drug response prediction performance when adding docking scores as additional features, through tests on large drug screening data. We discuss the limitations of the current approach and provide the research community with a baseline dataset of the large-scale computational docking for anti-cancer drugs.

60 APPLIED LIFE SCIENCES↗

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗