Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “molecular descriptors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Advanced Modeling and Process-Materials Co-Optimization Strategies for Swing Adsorption Based Gas Separations

This project devised a computational framework for simultaneously co-optimizing pressure swing adsorption process designs along with the sorbent materials (specifically, metal-organic frameworks) to be employed in the associated packed bed columns. The materials optimization aspect involved search over a design space that can describe the material’s molecular structure, while the process optimization aspect considered various process degrees of freedom for steps arising in various cycle configurations. This framework was demonstrated on the separation of nitrogen and carbon dioxide, which arises ubiquitously in a multitude of post-combustion carbon capture and “blue” hydrogen production applications. Our results led to metal-organic framework molecular descriptor choices that are predicted to outperform standard structures used in practice, providing guidance for future metal-organic framework synthesis efforts.

20 FOSSIL-FUELED POWER PLANTS↗

From Structured Solvents to Hybrid Materials (SS2HM) for Chemically Selective Capture and Electromagnetic Release of CO 2 : Mechanisms, Stability and Interfaces (Final Report)

The goal of this research program was to develop high capacity sorbents amenable for alternative regeneration approaches for direct air capture (DAC) of CO 2 . In particular, the research aimed to develop an understanding of CO 2 binding mechanism, thermal and oxidative stability, and regeneration energetics of functionalized ionic liquids (ILs), deep eutectic solvents (DESs), and porous materials. ILs and DESs are high-dielectric solvents with structural tunability that permits the rational-design for energy-efficient regeneration approaches based on electromagnetic (EM) field and moisture-swing. By further incorporating these solvents into polymeric capsules and other structural supports, multi-scale interfaces for targeted CO 2 and energy transfers were achieved. Aspects related to CO 2 capacity, selectivity, stability, dielectric properties, and binding energies were examined through experimental and computational design to identify molecular descriptors to inform future design of structured solvents and hybrid materials for DAC. Enclosed final report details the key findings, science advancements, and workforce development efforts from this project.

36 MATERIALS SCIENCE↗

Finch: Toxicity Dose Response Curve Prediction of Chemical Compounds and Mixtures

A paradigm shift in chemical risk assessment is emphasizing mixture testing over single compound analysis, eliminating animal testing, and adopting advanced modeling approaches to understand mixture activity profiles. However, existing computational models largely focus on single chemicals, with few effective solutions for modeling complex mixtures that account for synergistic or antagonistic effects and multiple Modes of Action (MoA). Conventional methods like concentration addition (CA) and independent action (IA) are insufficient for this task as they are designed for simplistic interactions and struggle to account for the dynamic and multifaceted nature of chemical mixtures, such as overlapping MoA and non-linear interactions. Finch offers a novel approach utilizing deep learning (DL) embeddings and multi-task quantitative structure-activity relationship (QSAR) models to improve chemical exposure prediction. By leveraging molecular descriptors, physiochemical properties, and large language model (LLM) embeddings from SMILES inputs, Finch preserves critical information in a latent space thereby enhancing predictive accuracy. The multi-task learning aspect of Finch is highly advantageous, as it simultaneously optimizes multiple loss functions, leveraging all available data across tasks to develop generalized representations that effectively capture complex ingredient interactions within mixtures.

59 BASIC BIOLOGICAL SCIENCES↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

Progress in Predicting Ionic Cocrystal Formation: The Case of Ammonium Nitrate

In contrast to the mature predictive frameworks applied to neutral cocrystals, ionic cocrystals, those including an ion pair, are difficult to design. Furthermore, they are generally excluded categorically from studies which correlate specific molecular properties to cocrystal formation, leaving the prospective ionic cocrystal engineer with few clear avenues to success. Herein ammonium nitrate, an energetic oxidizing salt, is targeted for cocrystallization in a potential coformer group selected based on likely interactions with the nitrate ion as revealed in the Cambridge Structural Database; six novel ionic cocrystals were discovered. Molecular descriptors previously identified as being related to neutral cocrystal formation were examined across the screening group but showed no relationship with ionic cocrystal formation. High packing coefficient is shown to be a constant among the successful coformers in the set and is utilized to directly target two more successful coformers, bypassing the need for a large screening group.

Andrew J. Bennett↗

E min : A First-Principles Thermochemical Descriptor for Predicting Molecular Synthesizability

Predicting the synthesizability of a new molecule remains an unsolved challenge that chemists have long tackled with heuristic approaches. Here, in this study, we report a new method for predicting synthesizability using a simple yet accurate thermochemical descriptor. We introduce E min , the energy difference between a molecule and its lowest energy constitutional isomer, as a synthesizability predictor that is accurate, physically meaningful, and first-principles based. We apply E min to 134,000 molecules in the QM9 data set and find that E min is accurate when used alone and reduces incorrect predictions of "synthesizable" by up to 52% when used to augment commonly used prediction methods. Our work illustrates how first-principles thermochemistry and heuristic approximations for molecular stability are complementary, opening a new direction for synthesizability prediction methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning for Prediction of Thermodynamic Descriptors

Our objective is to apply machine learning (ML) algorithms for the prediction of molecular catalysis descriptors from geometric properties derived from experimental crystallographic databases. Catalysis is often considered a “low-data” discipline that is poorly suited for ML methods. An exception is the extensive structural information that is available for molecular catalysts through the Cambridge Structural Database (CSD), which contains atomically precise molecular structures from X-ray diffraction analysis for >600K metal complexes. As a proof-of-principle, we targeted the prediction of hydricity, a thermodynamic property that provides understanding and control of catalytic hydride transfer. We built a training set composed of ~100 molecular complexes with a known hydricity and structural information from the CSD. This data set was converted into a machine-readable format using the smooth overlap of atomic positions (SOAP) representation and further labeled with simple electronic descriptors for the metal centers. Multiple different neural networks were trained on this data set, and the accuracy of the hydricity predictions ranged from < 2 kcal/mol to 20 kcal/mol. The accuracy of each model was highly sensitive to which compounds were in the train versus test set, underscoring the challenges associated with small and chemically diverse data sets. Finally, to further augment the data set, we attempted to experimentally measure several new hydricity values, however these experiments were unsuccessful due to undesired chemical reactivity of the selected complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integration of Computational Docking into Anti-Cancer Drug Response Prediction Models

Cancer is a heterogeneous disease in that tumors of the same histology type can respond differently to a treatment. Anti-cancer drug response prediction is of paramount importance for both drug development and patient treatment design. Although various computational methods and data have been used to develop drug response prediction models, it remains a challenging problem due to the complexities of cancer mechanisms and cancer-drug interactions. To better characterize the interaction between cancer and drugs, we investigate the feasibility of integrating computationally derived features of molecular mechanisms of action into prediction models. Specifically, we add docking scores of drug molecules and target proteins in combination with cancer gene expressions and molecular drug descriptors for building response models. The results demonstrate a marginal improvement in drug response prediction performance when adding docking scores as additional features, through tests on large drug screening data. We discuss the limitations of the current approach and provide the research community with a baseline dataset of the large-scale computational docking for anti-cancer drugs.

60 APPLIED LIFE SCIENCES↗

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Anion-Assisted Delivery of Multivalent Cations to Inert Electrodes

To understand and control key electrochemical processes - metal plating, corrosion, intercalation, etc. requires molecular-scale details of the active species at electrochemical interfaces and their mechanisms for de-solvation from the electrolyte. Using free energy sampling techniques we reveal the interfacial speciation of divalent cations in ether-based electrolytes and mechanisms for their delivery to an inert graphene electrode interface. Surprisingly, we find that anion solvophobicity drives a high population of anion-containing species to the interface that facilitate the delivery of divalent cations, even to negatively charged electrodes. Our simulations indicate that cation desolvation is greatly facilitated by cation-anion coupling. We propose anion solvophobicity as a molecular-level descriptor for rational design of electrolytes with increased efficiency for electrochemical processes limited by multivalent cation desolvation.

Electrochemical interfaces↗

SOMAS: a platform for data-driven material discovery in redox flow battery development

Abstract Aqueous organic redox flow batteries offer an environmentally benign, tunable, and safe route to large-scale energy storage. The energy density is one of the key performance parameters of organic redox flow batteries, which critically depends on the solubility of the redox-active molecule in water. Prediction of aqueous solubility remains a challenge in chemistry. Recently, machine learning models have been developed for molecular properties prediction in chemistry and material science. The fidelity of a machine learning model critically depends on the diversity, accuracy, and abundancy of the training datasets. We build a comprehensive open access organic molecular database “Solubility of Organic Molecules in Aqueous Solution” (SOMAS) containing about 12,000 molecules that covers wider chemical and solubility regimes suitable for aqueous organic redox flow battery development efforts. In addition to experimental solubility, we also provide eight distinctive quantum descriptors including optimized geometry derived from high-throughput density functional theory calculations along with six molecular descriptors for each molecule. SOMAS builds a critical foundation for future efforts in artificial intelligence-based solubility prediction models.

25 ENERGY STORAGE↗

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)↗

Prediction of impact sensitivity, heat of formation and heat of explosion using atomic connectivity

In these proceedings we revisit a large collection of explosives and explosive descriptors with the goal of predicting impact sensitivity using only local atomic environments that can be deciphered from molecular SMILES strings as descriptors without utilizing empirically measured values or computationally expensive electronic structure calculations. From the original database of nearly 500 descriptors, removing empirically measured and electronic structure values decreased the number of descriptors to 135, which we reduced to 18 the most important descriptors using Random Forests. The condensed model predicted impact sensitivity with essentially the same accuracy as the existing, more complex model (R 2 = 0.788 and RSME = 0.312), while remaining applicable to all types of explosives (Peroxides, azides, C-Nitros, Nitroamines, Nitrate Esters, etc.). In addition to impact sensitivity, we proposed similar models to accurately predict values heat of formation (ΔH f ) and heat of explosion (Q), with R 2 = 0.966 and 0.916, respectively. In conclusion, the work in these proceedings allows for prediction of explosive performance and sensitivity with only chemical structure information and an estimate of density.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Comprehensive Machine Learning Model for Metal–Ligand Binding Prediction: Applications in Chemistry and Biology

A machine-learning (ML) model that predicts metal–ligand binding constants was developed using the open-source Chemprop software. The model was trained on over 30,000 experimental log K 1 values, which include both protonation and metal–ligand stability constants, comprising over 3500 ligands and 10 2 metal ions from 73 total elements, thus generalizing beyond existing limited approaches, which focus only on specific metals or ligand families. The best-performing model included a combination of SMILES-based molecular representations along with descriptors for the metal ion and experimental conditions. It had an external test R 2 value of 0.942, and MAE value of 0.834. A “SMILES-only” simpler version also produced accurate predictions and preserved the binding trends, serving as a quick and easily accessible alternative for users without computational expertise. The SMILES-only model performed comparably to density functional theory (DFT) calculations but utilized a fraction of the computational resources. The model was successfully applied across diverse domains, including bioinorganic chemistry, heavy metal remediation, and sensor development and demonstrated its effectiveness as a rapid and reliable screening tool for both academic and industrial uses.

Ligands↗

Automated AI-driven Molecular Design for Therapeutic Discovery

In recent years, artificial intelligence and machine learning (AI/ML) approaches have revolutionized the process of designing new therapeutics, enabling scientists to rapidly respond to emerging threats from various pathogens. A prime example is the SARS-CoV-2 main protease, a key target for the development of antiviral inhibitors. In this study, we employed a novel, integrated approach that combines AI-driven iterative design of inhibitor candidates, screening based on physio-chemical properties and toxicity, physics-based computational modeling of protein-inhibitor interactions, and AI-assisted analysis of Native MS biophysical assay and characterization of designed candidates. Our deep learning 3D-scaffold model, which uses an input scaffold as a starting point, generated tens of thousands of compounds while preserving the key scaffold. To optimize these candidates, we calculated a comprehensive set of 136 descriptors, including both 2D and 3D molecular features, for compounds targeting the SARS-CoV-2 Main protease (Mpro) and a neurodegenerative disease-associated protein, cyclophilin (Cyp). The generated compounds were initially filtered based on their properties and then ranked according to their predicted binding affinity using our automated modeling and ML methods. Experimental validation of the Mpro candidates showing inhibitory activity demonstrates that our workflow can expedite the therapeutic discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗