Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “QSAR”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

Development of QSAR models to predict blood-brain barrier permeability

Assessing drug permeability across the blood-brain barrier (BBB) is important when evaluating the abuse potential of new pharmaceuticals as well as developing novel therapeutics that target central nervous system disorders. One of the gold-standard in vivo methods for determining BBB permeability is rodent log BB; however, like most in vivo methods, it is time-consuming and expensive. In the present study, two statistical-based quantitative structure-activity relationship (QSAR) models were developed to predict BBB permeability of drugs based on their chemical structure. The in vivo BBB permeability data were harvested for 921 compounds from publicly available literature, non-proprietary drug approval packages, and University of Washington’s Drug Interaction Database. The cross-validation performance statistics for the BBB models ranged from 82 to 85% in sensitivity and 80–83% in negative predictivity. Additionally, the performance of newly developed models was assessed using an external validation set comprised of 83 chemicals. Overall, performance of individual models ranged from 70 to 75% in sensitivity, 70–72% in negative predictivity, and 78–86% in coverage. The predictive performance was further improved to 93% in coverage by combining predictions across the two software programs. These new models can be rapidly deployed to predict blood brain barrier permeability of pharmaceutical candidates and reduce the use of experimental animals.

59 BASIC BIOLOGICAL SCIENCES↗

Model Choice Metrics to Optimize Profile-QSAR Performance

Predicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. Here, this method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the model’s ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement.

Biological and medical sciences, Computer science↗

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)↗

Chemical Recommender System: Replacement Suggestions for Small Molecules

The Chemical Recommender System (CRS) is an open-source, high-performance toolkit that enables real-time similarity searches across the complete PubChem database (over 50 million molecules) using commodity hardware. The CRS addresses critical limitations in existing chemical informatics platforms through a novel vector database infrastructure, extensible model integration capabilities, and complete algorithmic transparency. The system implements a vector database deployment with partitioned indexing that achieves a ~60x speedup over traditional approaches. A containerized model integration framework allows researchers to seamlessly incorporate custom predictive models into the full-scale search and scoring pipeline, while complete configurability of search parameters, filtering logic, and scoring functions provides capabilities not available in existing black-box solutions. Beyond structural similarity, the CRS integrates OPERA QSAR models for thermophysical and toxicity predictions, RDKit synthetic accessibility scoring, and user-defined models to compute weighted final replacement scores. The complete system is accessible through an interactive web application supporting real-time progress monitoring, post-processing score re-weighting, automated PDF reporting, and batch processing capabilities.

Nair, Parthiv Anand [Sandia National Laboratories ↗

Toxicology and Biodegradability of Blendstocks for Mixing Controlled Compression Ignition Combustion: Literature Review of Available Data

The goal of the U.S. Department of Energy’s Co-Optima project is to develop an understanding of potential biomass-derived blendstocks for use in the transportation sector. This report documents the estimated environmental and toxicological impact of potential bioblendstocks for mixing controlled compression ignition (MCCI) engines. The candidate molecules were: 2-nonanol; n-undecane; 2,6,10-trimethyldodecane; 5-ethyl-4-propylnonane; hexylhexanoate; methyldecanoate; dibutoxymethane; 4-butoxyheptane; dipentylether; renewable diesel; and soy biodiesel. To provide a robust comparison, diesel surrogate molecules were also considered: α-methylnaphthalene; decahydronaphthalene; 2,2,4,4,6,8,8-heptamethylnonane; n-butylcyclohexane; n-hexadecane; tetralin; and n-dodecylbenzene. The intent of this work was not to provide absolute answers and remove any molecules from consideration, but to provide more robust data for each molecule that can be used in subsequent evaluations. Little literature data was available for many of the molecules, both potential blendstocks and diesel surrogates, so a quantitative structural activity relationship (QSAR) approach was used to provide input into the relevant models. This assessment included an evaluation of compartment partitioning in the event of an accidental chemical spill, fate and transport indicators, estimated biodegradability, and human and environmental toxicology. The higher molecular weight and longer hydrocarbon chain lengths of these MCCI molecules reduce water solubility and mobility. Generally, the oxygenated molecules improved biodegradability metrics. The hydrocarbons farnesane and 5-Et-4-PrNonane are not predicted to be biodegradable. Human acute toxicity was found to be slightly to non-toxic for the MCCI blendstocks and the diesel surrogates. 2-Nonanol, undecane, MeDecanoate, dibutoxymethane, and dipentylether were likely developmental toxicants. Although poorly water soluble, all the compounds, whether Co-Optima blendstocks or diesel surrogates, showed significant toxicity to D. magna. The oxygenated blendstocks were less likely to bioaccumulate than the non-oxygenated blendstocks, which behaved similar to the diesel surrogates. Overall, this predicted analysis showed that the MCCI blendstocks are very similar to their diesel counterparts. No significant showstoppers were found that would indicate concern moving forward with additional research into these potential fuels. The results presented were based solely on predictive models and serve to provide information going forward for continued evaluation of these molecules.

09 BIOMASS FUELS↗

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Finch: Toxicity Dose Response Curve Prediction of Chemical Compounds and Mixtures

A paradigm shift in chemical risk assessment is emphasizing mixture testing over single compound analysis, eliminating animal testing, and adopting advanced modeling approaches to understand mixture activity profiles. However, existing computational models largely focus on single chemicals, with few effective solutions for modeling complex mixtures that account for synergistic or antagonistic effects and multiple Modes of Action (MoA). Conventional methods like concentration addition (CA) and independent action (IA) are insufficient for this task as they are designed for simplistic interactions and struggle to account for the dynamic and multifaceted nature of chemical mixtures, such as overlapping MoA and non-linear interactions. Finch offers a novel approach utilizing deep learning (DL) embeddings and multi-task quantitative structure-activity relationship (QSAR) models to improve chemical exposure prediction. By leveraging molecular descriptors, physiochemical properties, and large language model (LLM) embeddings from SMILES inputs, Finch preserves critical information in a latent space thereby enhancing predictive accuracy. The multi-task learning aspect of Finch is highly advantageous, as it simultaneously optimizes multiple loss functions, leveraging all available data across tasks to develop generalized representations that effectively capture complex ingredient interactions within mixtures.

59 BASIC BIOLOGICAL SCIENCES↗

CATMoS: Collaborative Acute Toxicity Modeling Suite

Background: Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. Objectives: The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different endpoints: LD50 value, U.S. Environmental Protection Agency hazard categories, Globally Harmonized System for Classification and Labelling hazard categories, very toxic chemicals (LD50 =50 mg/kg), and non-toxic chemicals (LD50 >2000 mg/kg). Methods: An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. Results: The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared to in vivo results. Discussion: CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for over 800,000 chemicals have been made available via the NTP’s Integrated Chemical Environment. The models are also implemented in a free, standalone open-source tool, OPERA, which allows predictions of new and untested chemicals to be made.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗