Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “molecular descriptors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Anion-Assisted Delivery of Multivalent Cations to Inert Electrodes

To understand and control key electrochemical processes - metal plating, corrosion, intercalation, etc. requires molecular-scale details of the active species at electrochemical interfaces and their mechanisms for de-solvation from the electrolyte. Using free energy sampling techniques we reveal the interfacial speciation of divalent cations in ether-based electrolytes and mechanisms for their delivery to an inert graphene electrode interface. Surprisingly, we find that anion solvophobicity drives a high population of anion-containing species to the interface that facilitate the delivery of divalent cations, even to negatively charged electrodes. Our simulations indicate that cation desolvation is greatly facilitated by cation-anion coupling. We propose anion solvophobicity as a molecular-level descriptor for rational design of electrolytes with increased efficiency for electrochemical processes limited by multivalent cation desolvation.

Electrochemical interfaces↗

SOMAS: a platform for data-driven material discovery in redox flow battery development

Abstract Aqueous organic redox flow batteries offer an environmentally benign, tunable, and safe route to large-scale energy storage. The energy density is one of the key performance parameters of organic redox flow batteries, which critically depends on the solubility of the redox-active molecule in water. Prediction of aqueous solubility remains a challenge in chemistry. Recently, machine learning models have been developed for molecular properties prediction in chemistry and material science. The fidelity of a machine learning model critically depends on the diversity, accuracy, and abundancy of the training datasets. We build a comprehensive open access organic molecular database “Solubility of Organic Molecules in Aqueous Solution” (SOMAS) containing about 12,000 molecules that covers wider chemical and solubility regimes suitable for aqueous organic redox flow battery development efforts. In addition to experimental solubility, we also provide eight distinctive quantum descriptors including optimized geometry derived from high-throughput density functional theory calculations along with six molecular descriptors for each molecule. SOMAS builds a critical foundation for future efforts in artificial intelligence-based solubility prediction models.

25 ENERGY STORAGE↗

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)↗

Prediction of impact sensitivity, heat of formation and heat of explosion using atomic connectivity

In these proceedings we revisit a large collection of explosives and explosive descriptors with the goal of predicting impact sensitivity using only local atomic environments that can be deciphered from molecular SMILES strings as descriptors without utilizing empirically measured values or computationally expensive electronic structure calculations. From the original database of nearly 500 descriptors, removing empirically measured and electronic structure values decreased the number of descriptors to 135, which we reduced to 18 the most important descriptors using Random Forests. The condensed model predicted impact sensitivity with essentially the same accuracy as the existing, more complex model (R 2 = 0.788 and RSME = 0.312), while remaining applicable to all types of explosives (Peroxides, azides, C-Nitros, Nitroamines, Nitrate Esters, etc.). In addition to impact sensitivity, we proposed similar models to accurately predict values heat of formation (ΔH f ) and heat of explosion (Q), with R 2 = 0.966 and 0.916, respectively. In conclusion, the work in these proceedings allows for prediction of explosive performance and sensitivity with only chemical structure information and an estimate of density.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A Comprehensive Machine Learning Model for Metal–Ligand Binding Prediction: Applications in Chemistry and Biology

A machine-learning (ML) model that predicts metal–ligand binding constants was developed using the open-source Chemprop software. The model was trained on over 30,000 experimental log K 1 values, which include both protonation and metal–ligand stability constants, comprising over 3500 ligands and 10 2 metal ions from 73 total elements, thus generalizing beyond existing limited approaches, which focus only on specific metals or ligand families. The best-performing model included a combination of SMILES-based molecular representations along with descriptors for the metal ion and experimental conditions. It had an external test R 2 value of 0.942, and MAE value of 0.834. A “SMILES-only” simpler version also produced accurate predictions and preserved the binding trends, serving as a quick and easily accessible alternative for users without computational expertise. The SMILES-only model performed comparably to density functional theory (DFT) calculations but utilized a fraction of the computational resources. The model was successfully applied across diverse domains, including bioinorganic chemistry, heavy metal remediation, and sensor development and demonstrated its effectiveness as a rapid and reliable screening tool for both academic and industrial uses.

Ligands↗

Bias free multiobjective active learning for materials design and discovery

The design rules for materials are clear for applications with a single objective. For most applications, however, there are often multiple, sometimes competing objectives where there is no single best material and the design rules change to finding the set of Pareto optimal materials. In this work, we leverage an active learning algorithm that directly uses the Pareto dominance relation to compute the set of Pareto optimal materials with desirable accuracy. We apply our algorithm to de novo polymer design with a prohibitively large search space. Using molecular simulations, we compute key descriptors for dispersant applications and drastically reduce the number of materials that need to be evaluated to reconstruct the Pareto front with a desired confidence. This work showcases how simulation and machine learning techniques can be coupled to discover materials within a design space that would be intractable using conventional screening approaches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated AI-driven Molecular Design for Therapeutic Discovery

In recent years, artificial intelligence and machine learning (AI/ML) approaches have revolutionized the process of designing new therapeutics, enabling scientists to rapidly respond to emerging threats from various pathogens. A prime example is the SARS-CoV-2 main protease, a key target for the development of antiviral inhibitors. In this study, we employed a novel, integrated approach that combines AI-driven iterative design of inhibitor candidates, screening based on physio-chemical properties and toxicity, physics-based computational modeling of protein-inhibitor interactions, and AI-assisted analysis of Native MS biophysical assay and characterization of designed candidates. Our deep learning 3D-scaffold model, which uses an input scaffold as a starting point, generated tens of thousands of compounds while preserving the key scaffold. To optimize these candidates, we calculated a comprehensive set of 136 descriptors, including both 2D and 3D molecular features, for compounds targeting the SARS-CoV-2 Main protease (Mpro) and a neurodegenerative disease-associated protein, cyclophilin (Cyp). The generated compounds were initially filtered based on their properties and then ranked according to their predicted binding affinity using our automated modeling and ML methods. Experimental validation of the Mpro candidates showing inhibitory activity demonstrates that our workflow can expedite the therapeutic discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Screening and Discovery of Metal Compound Active Sites for Strong and Selective Adsorption of N 2 in Air

Photocatalytic nitrogen fixation has the potential to provide a greener route for producing nitrogen-based fertilizers under ambient conditions. Computational screening is a promising route to discover new materials for the nitrogen fixation process, but requires identifying “descriptors” that can be efficiently computed. In this work, we argue that selectivity toward the adsorption of molecular nitrogen and oxygen can act as a key descriptor. A catalyst that can selectively adsorb nitrogen and resist poisoning of oxygen and other molecules present in air has the potential to facilitate the nitrogen fixation process under ambient conditions. Here we provide a framework for active site screening based on multifidelity density functional theory (DFT) calculations for a range of metal oxides, oxyborides, and oxyphosphides. The screening methodology consists of initial low-fidelity fixed geometry calculations and a second screening in which more expensive geometry optimizations were performed. The approach identifies promising active sites on several TiO 2 polymorph surfaces and a VBO 4 surface, and the full nitrogen reduction pathway is studied with the BEEF-vdW and HSE06 functionals on two active sites. The findings suggest that metastable TiO 2 polymorphs may play a role in photocatalytic nitrogen fixation, and that VBO 4 may be an interesting material for further studies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MODTRAN3: An update and recent validations against airborne high resolution interferometer measurements

MODTRAN, the Moderate Resolution Atmospheric Radiance and Transmittance Model, encompasses all the capabilities of LOWTRAN 7, the widely used 20 cm(exp -1) resolution radiance code, but incorporates a much more sensitive molecular band model with 2 cm(exp -1) resolution. MODTRAN contains many important elements that other band model based radiative transfer codes do not incorporate. It shares with FASCODE: spherical geometry, single and multiple scattering default atmospheric profile descriptors (gases, aerosols, clouds, fogs, and rain), and molecular continua (H2O, CO2, O3, O2, N2). In addition, it can calculate the solar/lunar direct and scattered radiation. MODTRAN3 was released to the general public in November 1994. It has several important features that the previous version, MODTRAN2, does not have. Chloro-fluorocarbon (CFC) and related heavy molecules (whose spectroscopic properties first appear on the HITRAN92 data base as temperature-dependent cross sections) have been incorporated into pseudo-band models, with provision for using both default and user supplied profiles. The addition of SO2 and O2 in the UV, along with upgraded ozone Chappuis bands in the visible is also part of MODTRAN3. An improved multiple scattering algorithm, the DIScrete Ordinate Radiative Transfer (DISORT) has also been incorporated into MODTRAN3. MODTRAN is very fast: simple timing runs of MODTRAN3 vs. FASCOD3 show an improvement of more than a factor of 100 for a typical 500 cm(exp -1) spectral interval and comparable vertical layering. Speed is an important consideration in heating/cooling rates calculations, where a large number of radiative transfer calculations are needed. The MODTRAN3 used in this study is based on HITRAN92, but as mentioned, above, it will be upgraded to HITRAN94 upon its release at the end of 1994. MODTRAN has been adopted by: some researchers in the AVIRIS program as one radiative transfer code to derive surface reflectance from AVIRIS measurements. The accuracy of the code is very important because any errors in the radiative transfer calculation will directly translate into errors in the derived surface reflectance. In this paper, the new solar irradiance calculated by Kurucz, which is adopted in MODTRAN3, will be presented. Recent validations of MODTRAN3 with airborne high resolution interferometer measurements over ocean will be discussed. Good agreeement between model calculations and measurements was achieved.

Anderson, Gail P.↗

Thermodynamic and Kinetic Activity Descriptors for the Catalytic Hydrogenation of Ketones

Activity descriptors are a powerful tool for the design of catalysts than can efficiently utilize H 2 with minimal energy losses. In this study, we develop the use of hydricity and H - self-exchange rates as thermodynamic and kinetic descriptors for the hydrogenation of ketones by molecular catalysts. Two complexes with known hydricity, HRh(dmpe) 2 and HCo(dmpe) 2 , were investigated for the catalytic hydrogenation of ketones under mild conditions (1.5 atm, 25 °C). The rhodium catalyst proved to be an efficient catalyst for a wide range of ketones, whereas the cobalt catalyst could only hydrogenate electron-deficient ketones. Using a combination of experiment and electronic structure theory, thermodynamic hydricity values were established for 46 alkoxide/ketone pairs in both MeCN and THF solvent. Through comparison of the hydricities of the catalysts and substrates, it was determined that catalysis was only observed for catalyst/ketone pairs with an exergonic H - transfer step. Mechanistic studies revealed that H - transfer was rate-limiting step for catalysis, allowing for the experimental and computation construction of linear free-energy relationships (LFERs) for H - transfer. Further analysis revealed the LFERs could be reproduced using Marcus theory, in which the H - self-exchange rates for the HRh/Rh + and ketone/alkoxide pairs were used to predict the experimentally measured catalytic barriers within 2 kcal mol -1 . Finally, these studies significantly expand the scope of catalytic reactions that can be analyzed with a thermodynamic hydricity descriptor and firmly establish Marcus theory as a valid approach to develop kinetic descriptors for designing catalysts for H - transfer reactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining artificial intelligence and physics-based modeling to directly assess atomic site stabilities: from sub-nanometer clusters to extended surfaces

The performance of functional materials is dictated by chemical and structural properties of individual atomic sites. In catalysts, for instance, the thermodynamic stability of constituting atomic sites is a key descriptor from which more complex properties, such as molecular adsorption energies and reaction rates, can be derived. In this study, we present a widely applicable machine learning (ML) approach to instantaneously compute the stability of individual atomic sites in structurally and electronically complex nano-materials. Conventionally, we determine such site stabilities using computationally intensive first-principles calculations. With our approach, we predict the stability of atomic sites in sub-nanometer metal clusters of 3–55 atoms with mean absolute errors in the range of 0.11–0.14 eV. To extract physical insights from the ML model, we introduce a genetic algorithm (GA) for feature selection. This algorithm distills the key structural and chemical properties governing the stability of atomic sites in size-selected nanoparticles, allowing for physical interpretability of the models and revealing structure–property relationships. The results of the GA are generally model and materials specific. In the limit of large nanoparticles, the GA identifies features consistent with physics-based models for metal–metal interactions. By combining the ML model with the physics-based model, we predict atomic site stabilities in real time for structures ranging from sub-nanometer metal clusters (3–55 atom) to larger nanoparticles (147 to 309 atoms) to extended surfaces using a physically interpretable framework. Finally, we present a proof of principle showcasing how our approach can determine stable and active nanocatalysts across a generic materials space of structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning predictions of diffusion in bulk and confined ionic liquids using simple descriptors

Ionic liquids have many intriguing properties and widespread applications such as separations and energy storage. However, ionic liquids are complex fluids and predicting their behavior is difficult, particularly in confined environments. We introduce fast and computationally efficient machine learning (ML) models that can predict diffusion coefficients and ionic conductivity of bulk and nanoconfined ionic liquids over a wide temperature range (350–500 K). The ML models are trained on molecular dynamics simulation data for 29 unique ionic liquids as bulk fluids and confined in graphite slit pores. This model is based on simple physical descriptors of the cations and anions such as molecular weight and surface area. Here, we also demonstrate that accurate results can be obtained using only descriptors derived from SMILES (simplified molecular-input line-entry system) codes for the ions with minimal computational effort. This offers a fast and efficient method for estimating diffusion and conductivity of nanoconfined ionic liquids at various temperatures without the need for expensive molecular dynamics simulations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Equivariant Graph Attention Network - 3D Conformers & Feature Fusion

EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.

Shim, Heesung [Lawrence Livermore National Laborat↗

Identification of novel organic polar materials: A machine learning study with importance sampling

Recent advances in the synthesis of polar molecular materials have produced practical alternatives to ferroelectric ceramics, opening up exciting new avenues for their incorporation into modern electronic devices. However, in order to realize the full potential of polar polymer and molecular crystals for modern technological applications, it is paramount to assemble and evaluate all the available data for such compounds, identifying descriptors that could be associated with an emergence of ferroelectricity. In this paper, we utilized data-driven approaches to judiciously shortlist candidate materials from a wide chemical space that could possess ferroelectric functionalities. A machine learning study with importance sampling was employed to address the challenge of having a limited amount of available data on already-known organic ferroelectrics. Sets of molecular- and crystal-level descriptors were combined with a Random Forest Regression algorithm in order to predict the spontaneous polarization of the shortlisted compounds. First-principles simulations were performed to further validate the predictions obtained from the machine learning model.

36 MATERIALS SCIENCE↗

A study of real-world micrograph data quality and machine learning model robustness

Abstract Machine-learning (ML) techniques hold the potential of enabling efficient quantitative micrograph analysis, but the robustness of ML models with respect to real-world micrograph quality variations has not been carefully evaluated. We collected thousands of scanning electron microscopy (SEM) micrographs for molecular solid materials, in which image pixel intensities vary due to both the microstructure content and microscope instrument conditions. We then built ML models to predict the ultimate compressive strength (UCS) of consolidated molecular solids, by encoding micrographs with different image feature descriptors and training a random forest regressor, and by training an end-to-end deep-learning (DL) model. Results show that instrument-induced pixel intensity signals can affect ML model predictions in a consistently negative way. As a remedy, we explored intensity normalization techniques. It is seen that intensity normalization helps to improve micrograph data quality and ML model robustness, but microscope-induced intensity variations can be difficult to eliminate.

36 MATERIALS SCIENCE↗