Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “toxicity prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Spatiotemporal predictions of toxic urban plumes using deep learning

Industrial accidents, chemical spills, and structural fires can release large amounts of harmful materials that disperse into urban atmospheres and impact populated areas. Computer models are typically used to predict the transport of toxic plumes by solving fluid dynamical equations. However, these models can be computationally expensive due to the need for many grid cells to simulate turbulent flow and resolve individual buildings and streets. In emergency response situations, alternative methods are needed that can run quickly and adequately capture important spatiotemporal features. Here, we present a novel deep learning model called ST-GasNet inspired by the mathematical equations that govern the behavior of plumes as they disperse through the atmosphere. ST-GasNet learns the spatiotemporal dependencies from a limited set of temporal sequences of ground-level toxic urban plumes generated by a high-resolution large eddy simulation model. On independent sequences, ST-GasNet accurately predicts the late-time spatiotemporal evolution, given the early-time behavior as an input, even when a building splits a large plume into smaller plumes. By incorporating large-scale wind boundary condition information, ST-GasNet achieves a prediction accuracy of at least 90% on test data for the entire prediction period.

Civil and Environmental Engineering↗

Microbial vitamin biosynthesis links gut microbiota dynamics to chemotherapy toxicity

ABSTRACT Dose-limiting toxicities pose a major barrier to cancer treatment. While preclinical studies show that the gut microbiota influences and is influenced by anticancer drugs, data from patients paired with careful side effect monitoring remains limited. Here, we investigate capecitabine (CAP)-microbiome interactions through longitudinal metagenomic sequencing of stool from 56 advanced colorectal cancer patients. CAP significantly altered the gut microbiome, enriching for menaquinol (vitamin K2) biosynthesis genes. Transposon library screens, targeted gene deletions, and media supplementation revealed that menaquinol biosynthesis protectsEscherichia colifrom drug toxicity. Stool menaquinol gene and metabolite levels were associated with decreased peripheral sensory neuropathy. Machine learning models trained in this cohort predicted toxicities in an independent cohort. Taken together, these results suggest treatment-associated increases in microbial vitamin biosynthesis serve a chemoprotective role for bacterial and host cells. Further, our findings provide a foundation for in-depth mechanistic dissection, human intervention studies, and extension to other cancer treatments. IMPORTANCE Side effects are common during the treatment of cancer. The trillions of microbes found within the human gut are sensitive to anticancer drugs, but the effects of treatment-induced shifts in gut microbes for side effects remain poorly understood. We profiled gut microbes in colorectal cancer patients treated with capecitabine and carefully monitored side effects. We observed a marked expansion in genes for producing vitamin K2 (menaquinone). Vitamin K2 rescued gut bacterial growth and was associated with decreased side effects in patients. We then used information about gut microbes to develop a predictive model of drug toxicity that was validated in an independent cohort. These results suggest that treatment-associated increases in bacterial vitamin production protect both bacteria and host cells from drug toxicity, providing new opportunities for intervention and motivating the need to better understand how dietary intake and bacterial production of micronutrients like vitamin K2 influence cancer treatment outcomes.

Microbiology↗

Chemical Recommender System: Replacement Suggestions for Small Molecules

The Chemical Recommender System (CRS) is an open-source, high-performance toolkit that enables real-time similarity searches across the complete PubChem database (over 50 million molecules) using commodity hardware. The CRS addresses critical limitations in existing chemical informatics platforms through a novel vector database infrastructure, extensible model integration capabilities, and complete algorithmic transparency. The system implements a vector database deployment with partitioned indexing that achieves a ~60x speedup over traditional approaches. A containerized model integration framework allows researchers to seamlessly incorporate custom predictive models into the full-scale search and scoring pipeline, while complete configurability of search parameters, filtering logic, and scoring functions provides capabilities not available in existing black-box solutions. Beyond structural similarity, the CRS integrates OPERA QSAR models for thermophysical and toxicity predictions, RDKit synthetic accessibility scoring, and user-defined models to compute weighted final replacement scores. The complete system is accessible through an interactive web application supporting real-time progress monitoring, post-processing score re-weighting, automated PDF reporting, and batch processing capabilities.

Nair, Parthiv Anand [Sandia National Laboratories ↗

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

Finch: Toxicity Dose Response Curve Prediction of Chemical Compounds and Mixtures

A paradigm shift in chemical risk assessment is emphasizing mixture testing over single compound analysis, eliminating animal testing, and adopting advanced modeling approaches to understand mixture activity profiles. However, existing computational models largely focus on single chemicals, with few effective solutions for modeling complex mixtures that account for synergistic or antagonistic effects and multiple Modes of Action (MoA). Conventional methods like concentration addition (CA) and independent action (IA) are insufficient for this task as they are designed for simplistic interactions and struggle to account for the dynamic and multifaceted nature of chemical mixtures, such as overlapping MoA and non-linear interactions. Finch offers a novel approach utilizing deep learning (DL) embeddings and multi-task quantitative structure-activity relationship (QSAR) models to improve chemical exposure prediction. By leveraging molecular descriptors, physiochemical properties, and large language model (LLM) embeddings from SMILES inputs, Finch preserves critical information in a latent space thereby enhancing predictive accuracy. The multi-task learning aspect of Finch is highly advantageous, as it simultaneously optimizes multiple loss functions, leveraging all available data across tasks to develop generalized representations that effectively capture complex ingredient interactions within mixtures.

59 BASIC BIOLOGICAL SCIENCES↗

A mixture parameterized biologically based dosimetry model to predict body burdens of polycyclic aromatic hydrocarbons in developmental zebrafish toxicity assays

Polycyclic aromatic hydrocarbons (PAHs) are a group of environmental toxicants found ubiquitously as complex mixtures in human-impacted environments. Developmental zebrafish exposures have been used widely to study PAH toxicity, but most studies report nominal exposure concentrations. Nominal exposure concentrations can be unreliable dose metrics due to differences in toxicant bioavailability resulting from disparate exposure methodologies and chemical properties. Toxicokinetic modeling can predict toxicant tissue doses to facilitate comparison between exposures of different chemicals, methodologies, and biological models. We parameterize a biologically based dosimetry model for developmental zebrafish toxicity assays for 9 PAHs. The model was optimized with measurements from media, tissue, and plastic plate walls throughout a static developmental exposure to a mixture of 10 PAHs of high abundance within the Portland Harbor Superfund Site. Plate binding, volatilization, zebrafish permeability, and tissue—media partitioning coefficients vary widely between PAHs. Model predictions accounted for 83% and 54% of 48 hpf body burdens within a factor of 2 resulting from exposures to mixtures and individual PAHs, respectively. Accounting for solubility significantly improves model performance. Competition for active sites in metabolizing enzymes may change biotransformation kinetics between individual PAH and mixture exposures. Area under the curve estimations of concentrations in zebrafish resulted in altered hazard rankings from nominal exposure concentrations. Future work will be oriented to generalizing the model to other PAHs. This PAH dosimetry model improves the interpretability of developmental zebrafish toxicity assays by providing time-resolved body burdens from nominal exposure concentrations.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Conotoxin Prediction: New Features to Increase Prediction Accuracy

Conotoxins are toxic, disulfide-bond-rich peptides from cone snail venom that target a wide range of receptors and ion channels with multiple pathophysiological effects. Conotoxins have extraordinary potential for medical therapeutics that include cancer, microbial infections, epilepsy, autoimmune diseases, neurological conditions, and cardiovascular disorders. Despite the potential for these compounds in novel therapeutic treatment development, the process of identifying and characterizing the toxicities of conotoxins is difficult, costly, and time-consuming. This challenge requires a series of diverse, complex, and labor-intensive biological, toxicological, and analytical techniques for effective characterization. While recent attempts, using machine learning based solely on primary amino acid sequences to predict biological toxins (e.g., conotoxins and animal venoms), have improved toxin identification, these methods are limited due to peptide conformational flexibility and the high frequency of cysteines present in toxin sequences. This results in an enumerable set of disulfide-bridged foldamers with different conformations of the same primary amino acid sequence that affect function and toxicity levels. Consequently, a given peptide may be toxic when its cysteine residues form a particular disulfide-bond pattern, while alternative bonding patterns (isoforms) or its reduced form (free cysteines with no disulfide bridges) may have little or no toxicological effects. Similarly, the same disulfide-bond pattern may be possible for other peptide sequences and result in different conformations that all exhibit varying toxicities to the same receptor or to different receptors. We present here new features, when combined with primary sequence features to train machine learning algorithms to predict conotoxins, that significantly increase prediction accuracy.

collisional cross section↗

Building a predictive model for polycyclic aromatic hydrocarbon dosimetry in organotypically cultured human bronchial epithelial cells using benzo[ a ]pyrene

The airway epithelium is a primary route of exposure for inhaled toxicants, and organotypic culture models represent an important advancement for toxicity testing compared to simple in vitro models that may lack metabolic capability and multicellular structure/communication associated with the bronchial epithelium in vivo. A quantitative understanding of chemical dosimetry is key for interpreting and extrapolating study results; however, dosimetry is understudied in organotypic models limiting ability to predict toxicity. We developed a dosimetry model for primary human bronchial epithelial cells (HBECs) cultured at the air-liquid interface (ALI) using benzo[a]pyrene (BAP), a representative polycyclic aromatic hydrocarbon. Dose and time course evaluation of metabolite formation and enzyme activity and expression were utilized to parameterize a cellular dosimetry model to improve the utility of ALI-HBECs for assessing chemical risk. Dosimetry analysis demonstrated absorption of BAP into cells and an increase in Phase 1 and 2 metabolites over time that correlated with regulation of metabolizing enzymes. BAP was cleared from cells by 48 h after exposure, and the primary metabolites generated in ALI-HBECs were BAP-3-phenol, BAP-4,5-dihydrodiol, BAP-7,8-dihydrodiol, BAP-9,10-dihydrodiol, BAP-7,8,9,10-tetrol, BAP-3-phenol-glucuronide, BAP-4,5-dihydrodiol-glucuronide, and BAP-9,10-dihydrodiol-glucuronide. The resulting dosimetry model described BAP and 7,8-dihydrodiol toxicokinetics in ALI-HBECs and suggested active excretion of 7,8-dihydrodiol. Overall, this study demonstrates metabolic competency of ALI-HBECs for BAP metabolism, demonstrates the usefulness of complex in vitro systems for human-relevant toxicity data, and exhibits how in silico models can be utilized for understanding the dosimetry of test compounds to aid in in vitro to human extrapolation of toxicity data for risk assessments.

Benzo[a]pyrene↗

Machine Learning Framework for Conotoxin Class and Molecular Target Prediction

Conotoxins are small and highly potent neurotoxic peptides derived from the venom of marine cone snails which have captured the interest of the scientific community due to their pharmacological potential. These toxins display significant sequence and structure diversity, which results in a wide range of specificities for several different ion channels and receptors. Despite the recognized importance of these compounds, our ability to determine their binding targets and toxicities remains a significant challenge. Predicting the target receptors of conotoxins, based solely on their amino acid sequence, remains a challenge due to the intricate relationships between structure, function, target specificity, and the significant conformational heterogeneity observed in conotoxins with the same primary sequence. We have previously demonstrated that the inclusion of post-translational modifications, collisional cross sections values, and other structural features, when added to the standard primary sequence features, improves the prediction accuracy of conotoxins against non-toxic and other toxic peptides across varied datasets and several different commonly used machine learning classifiers. Here, we present the effects of these features on conotoxin class and molecular target predictions, in particular, predicting conotoxins that bind to nicotinic acetylcholine receptors (nAChRs). We also demonstrate the use of the Synthetic Minority Oversampling Technique (SMOTE)-Tomek in balancing the datasets while simultaneously making the different classes more distinct by reducing the number of ambiguous samples which nearly overlap between the classes. In predicting the alpha, mu, and omega conotoxin classes, the SMOTE-Tomek PCA PLR model, using the combination of the SS and P feature sets establishes the best performance with an overall accuracy (OA) of 95.95%, with an average accuracy (AA) of 93.04%, and an f1 score of 0.959. Using this model, we obtained sensitivities of 98.98%, 89.66%, and 90.48% when predicting alpha, mu, and omega conotoxin classes, respectively. Similarly, in predicting conotoxins that bind to nAChRs, the SMOTE-Tomek PCA SVM model, which used the collisional cross sections (CCSs) and the P feature sets, demonstrated the highest performance with 91.3% OA, 91.32% AA, and an f1 score of 0.9131. The sensitivity when predicting conotoxins that bind to nAChRs is 91.46% with a 91.18% sensitivity when predicting conotoxins that do not bind to nAChRs.

59 BASIC BIOLOGICAL SCIENCES↗

Connecting suborganismal data to bioenergetic processes: killifish embryos exposed to a dioxin-like compound

A core challenge for ecological risk assessment is to integrate molecular responses into a chain of causality to organismal or population level outcomes. Bioenergetic theory may be a useful approach for integrating suborganismal responses to predict organismal level responses that influence population dynamics. In this work, we describe a novel application of Dynamic Energy Budget (DEB), theory in the context of a toxicity framework (Adverse Outcome Pathways, AOP) to make quantitative predictions of chemical exposures to individuals, starting from suborganismal data. We use early life stage exposure of Fundulus heteroclitus to dioxin-like chemicals (DLCs) and connect AOP Key Events (KEs) to DEB processes through “damage” that is produced at a rate proportional to the internal toxicant concentration. We use transcriptomic data of fish embryos exposed to DLCs to translate molecular indicators of damage into changes in DEB parameters (damage increases somatic maintenance costs) and use DEB models to predict sublethal and lethal effects of young fish. By changing a small subset of model parameters, we predict the evolved tolerance to DLCs in some wild F. heteroclitus populations, a data set not used in model parameterization. The differences in model parameters points to reduced sensitivity and altered damage repair dynamics as contributing to this evolved resistance. Our methodology has potential extrapolation to untested chemicals of ecological concern.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Aryl hydrocarbon receptor-dependent toxicity by retene requires metabolic competence

Polycyclic aromatic hydrocarbons (PAHs) are a class of organic compounds frequently detected in the environment with widely varying toxicities. Many PAHs activate the aryl hydrocarbon receptor (AHR), inducing the expression of a battery of genes, including xenobiotic metabolizing enzymes like cytochrome P450s (CYPs); however, not all PAHs act via this mechanism. We screened several parent and substituted PAHs in in vitro AHR activation assays to classify their unique activity. Retene (1-methyl-7-isopropylphenanthrene) displays Ahr2-dependent teratogenicity in zebrafish, but did not activate human AHR or zebrafish Ahr2, suggesting a retene metabolite activates Ahr2 in zebrafish to induce developmental toxicity. To investigate the role of metabolism in retene toxicity, studies were performed to determine the functional role of cyp1a, cyp1b1, and the microbiome in retene toxicity, identify the zebrafish window of susceptibility, and measure retene uptake, loss, and metabolite formation in vivo. Cyp1a-null fish were generated using CRISPR-Cas9. Cyp1a-null fish showed increased sensitivity to retene toxicity, whereas Cyp1b1-null fish were less susceptible, and microbiome elimination had no significant effect. Zebrafish required exposure to retene between 24 and 48 hours post fertilization (hpf) to exhibit toxicity. After static exposure, retene concentrations in zebrafish embryos increased until 24 hpf, peaked between 24 and 36 hpf, and decreased rapidly thereafter. We detected retene metabolites at 36 and 48 hpf, indicating metabolic onset preceding toxicity. This study highlights the value of combining molecular and systems biology approaches with mechanistic and predictive toxicology to interrogate the role of biotransformation in AHR-dependent toxicity.

59 BASIC BIOLOGICAL SCIENCES↗

Methods for cytotoxic chemotherapy-based predictive assays

The invention relates to methods, systems and kits for determining therapeutic effectiveness or toxicity of cancer-treating compounds that incorporate into or bind to DNA. In particular, the invention is directed to methods, systems and kits for predicting a patient's treatment outcome after administration of a microdose of therapeutic composition to the patient. The methods provides physicians with a diagnostic tool to segregate cancer patients into differential populations that have a higher or lower chance of responding to a particular therapeutic treatment.

Henderson, Paul↗

Transcriptional pathways linked to fetal and maternal hepatic dysfunction caused by gestational exposure to perfluorooctanoic acid (PFOA) or hexafluoropropylene oxide-dimer acid (HFPO-DA or GenX) in CD-1 mice

Per- and polyfluoroalkyl substances (PFAS) comprise a diverse class of chemicals used in industrial processes, consumer products, and fire-fighting foams which have become environmental pollutants of concern due to their persistence, ubiquity, and associations with adverse human health outcomes, including in pregnant persons and their offspring. Multiple PFAS are associated with adverse liver outcomes in adult humans and toxicological models, but effects on the developing liver are not fully described. Here we performed transcriptomic analyses in the mouse to investigate the molecular mechanisms of hepatic toxicity in the dam and its fetus after exposure to two different PFAS, perfluorooctanoic acid (PFOA) and its replacement, hexafluoropropylene oxide-dimer acid (HFPO-DA, known as GenX). Pregnant CD-1 mice were exposed via oral gavage from embryonic day (E) 1.5-17.5 to PFOA (0, 1, or 5 mg/kg-d) or GenX (0, 2, or 10 mg/kg-d). Maternal and fetal liver RNA was isolated (N = 5 per dose/group) and the transcriptome analyzed by Affymetrix Array. Differentially expressed genes (DEG) and differentially enriched pathways (DEP) were obtained. DEG patterns were similar in maternal liver for 5 mg/kg PFOA, 2 mg/kg GenX, and 10 mg/kg GenX (R2: 0.46-0.66). DEG patterns were similar across all 4 dose groups in fetal liver (R2: 0.59-0.81). There were more DEGs in fetal liver compared to maternal liver at the low doses for both PFOA (fetal = 69, maternal = 8) and GenX (fetal = 154, maternal = 93). Upregulated DEPs identified across all groups included Fatty Acid Metabolism, Peroxisome, Oxidative Phosphorylation, Adipogenesis, and Bile Acid Metabolism. Transcriptome-phenotype correlation analyses demonstrated > 1000 maternal liver DEGs were significantly correlated with maternal relative liver weight (R 2 >0.92). These findings show shared biological pathways of liver toxicity for PFOA and GenX in maternal and fetal livers in CD-1 mice. The limited overlap in specific DEGs between the dam and fetus suggests the developing liver responds differently than the adult liver to these chemical stressors. This work helps define mechanisms of hepatic toxicity of two structurally unique PFAS and may help predict latent consequences of developmental exposure.

54 ENVIRONMENTAL SCIENCES↗

Enhanced Hanford High-Fluoride Waste Glass Property Data Development: Phase 1

This study focused on investigating the effects of fluorine concentration on simulated high-level waste glass properties to eventually establish a fluorine limit (as a single-component or multiple-component constraint) for glass formulations for high-fluoride Hanford wastes. This is a first step to provide data to understand the impacts of changing flowsheets on the mission duration and extent. A test matrix of 20 high-fluoride glasses was generated, and the chemical compositions were measured. The following properties were measured and tested against current model predictions: crystal formation after centerline canister cooling, crystallinity as a function of temperature, density, viscosity, electrical conductivity, toxic leaching characteristics using the toxicity characteristic leach profile (TCLP), product consistency using the product consistency test (PCT), and SO 3 solubility. Overall, current models failed to adequately predict most of the properties, possibly due to differences in compositional space used to generate the models and the current test matrix. Additional work is needed to more accurately assess the impacts of high-fluoride wastes on Hanford processing, including additional data collection over a broader composition region and model development for the key models of interest such as PCT and TCLP.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗