SEARCH · Engineering Papers
Results for “computational molecular design”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction
Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.
End-to-end optimization for battery materials and molecules by combining graph neural networks and reinforcement learning
The National Renewable Energy Laboratory (NREL), together with the Colorado School of Mines (CSM) and Colorado State University (CSU), has developed a machine learning-enhanced approach to design new battery materials. Currently, such materials are designed in part via numerous expensive high-fidelity computational simulations that predict the performance of a given composition. Even with computational screening tools, the vast landscape of possible molecular or crystal structures exceeds current and future computational capacity. Improving the efficiency by which new materials can be optimized will therefore disrupt the cost, risk, and time required to bring new energy solutions to the marketplace. Predicting the properties of an organic molecule or periodic crystalline material given its structure has grown increasingly common. These approaches leverage large-scale computational and experimental databases and ML approaches such as graph neural networks. The inverse design problem of finding a material that possesses desired properties is substantially more challenging, since enumerating all valid material structures is not feasible. In this project, we leveraged recent success in reinforcement learning to efficiently navigate this high-dimensional search space. Just as algorithms can find the optimal chess moves from nearly limitless options, we train an approach to evolve a simple starting structure into a complex structure that possess the desired properties. Our solution has been demonstrated by applying it to two related design application tasks for short- and long-term energy storage, respectively: (1) the design of solid-state ion conductors and (2) the design of organic redox-active materials. The project has resulted an open-source software library for material design, documented examples of applying the library to both organic and inorganic material optimization, and peer-reviewed publications detailing the data, computational models, and resulting candidate materials.
Using Ultrafast Entangled Photon Correlations to Measure the Temporal Evolution of Optically Excited Molecular Entanglement (Final Technical Report)
The goal of this project is to build an ultrafast entangled photon spectrometer and use it to measure theoretical predictions that photoexcited states are enhanced by spin-photon interactions. Entanglement of two states describes a specific type of quantum superposition in which measuring one state gives information about a second state. While entanglement is well explored for quantum information and computing systems, its effects on photoexcited states are less understood, especially in the ultrafast domains of molecular vibronic coupling. The grant designs a frequency and temporally resolved entangled photon spectrometer and uses it to test theoretical predictions like enhanced two photon absorption, non-reciprocal Fourier relations, and coupling of photons to spin systems. The end goal of the project is an understanding of how, when, and where entanglement is useful for spectroscopy.
Actinide Molten Salts: A Machine-Learning Potential Molecular Dynamics Study
We know that actinide molten salts represent a class of important materials in nuclear energy. Understanding them at a molecular level is critical to proper and optimal design of relevant technological applications. Yet, owing to the complexity of electronic structure due to the 5f orbitals, computational studies of heavy elements in condensed phases using ab initio potentials to study the structure and dynamics of these elements embedded in molten salts are difficult. This lack of efficient computational protocols makes it difficult to obtain information on properties that require extensive statistical sampling like transport. To tackle this problem, we adopted a machine-learning approach to study ThCl 4 -NaCl and UCl 3 -NaCl binary systems. The machine-learning potential, with the density functional theory accuracy, allows us to obtain long molecular dynamics trajectories (ns) for large systems (10 3 atoms) at a considerably low computing cost, thereby efficiently gaining information about their bonding structures, thermodynamics, and dynamics at a range of temperature. We observed a considerable change in the coordination environments of actinide elements and their characteristic coordination-sphere lifetime. Our study also suggests that actinides in molten salts may not follow well known entropy-scaling laws.
PySpawn: Software for Nonadiabatic Quantum Molecular Dynamics
The ab initio multiple spawning (AIMS) method enables nonadiabatic quantum molecular dynamics simulations in an arbitrary number of dimensions, with potential energy surfaces provided by electronic structure calculations performed on-the-fly. However, the intricacy of the AIMS algorithm complicates software development, deployment on modern shared computer resources, and post-simulation data analysis. PySpawn is a nonadiabatic molecular dynamics software package that addresses these issues. Here, the program is designed to be easily interfaced with electronic structure software, and an interface to the TeraChem software package is described here. PySpawn introduces a task-based reorganization of the AIMS algorithm, allowing fine-grained restart capability and setting the stage for efficient parallelization in a future release. PySpawn includes a user-friendly and interactive Python analysis module that will enable novice users to painlessly adopt AIMS. As a demonstration of PySpawn’s simulation capability and analysis module, we report complete active space self-consistent field–based AIMS simulations of the 1,2- dithienyl-1,2-dicyanoethene molecule, a promising molecular photoswitch.
Efficient Quantum Gibbs Samplers with Kubo–Martin–Schwinger Detailed Balance Condition
Lindblad dynamics and other open-system dynamics provide a promising path towards efficient Gibbs sampling on quantum computers. In these proposals, the Lindbladian is obtained via an algorithmic construction akin to designing an artificial thermostat in classical Monte Carlo or molecular dynamics methods, rather than being treated as an approximation to weakly coupled system-bath unitary dynamics. Recently, Chen, Kastoryano, and Gilyén (arXiv:2311.09207) introduced the first efficiently implementable Lindbladian satisfying the Kubo–Martin–Schwinger (KMS) detailed balance condition, which ensures that the Gibbs state is a fixed point of the dynamics and is applicable to non-commuting Hamiltonians. This Gibbs sampler uses a continuously parameterized set of jump operators, and the energy resolution required for implementing each jump operator depends only logarithmically on the precision and the mixing time. In this work, we build upon the structural characterization of KMS detailed balanced Lindbladians by Fagnola and Umanità, and develop a family of efficient quantum Gibbs samplers using a finite set of jump operators (the number can be as few as one), akin to the classical Markov chain-based sampling algorithm. Compared to the existing works, our quantum Gibbs samplers have a comparable quantum simulation cost but with greater design flexibility and a much simpler implementation and error analysis. Moreover, it encompasses the construction of Chen, Kastoryano, and Gilyén as a special instance.
Multigene engineering in plants: Technologies, applications, and future prospects
The emerging bioeconomy presents a promising solution to both economic and environmental challenges. Within the bioeconomy, plants serve as a renewable, sustainable, and cost-effective source of foods, fuels, chemicals, and materials. However, traditional breeding and single-gene engineering approaches fall short in addressing complex traits (e.g., drought tolerance, disease resistance, yield, nutrient use efficiency) which are controlled by multiple genes. The complexity of plant biology often necessitates the use of multigene engineering (MGE), which involves simultaneous ectopic expression, up/down-regulation, or editing of multiple genes, to enhance plant traits relevant to the bioeconomy. These genes may be associated with distinct traits or function as components of specific metabolic and regulatory pathways. This review summarizes current technologies for MGE within the synthetic biology-driven Design-Build-Test-Learn (DBTL) framework, detailing its four key stages: Design – gene construct development; Build – DNA assembly and plant transformation; Test – the molecular, biochemical, and physiological characterization of engineered plants; and Learn – computational modeling to refine, multiplex and iterate the process. Despite good progress in the applications of MGE in biofortification, metabolic engineering, and stress resilience, challenges remain in construct stability, coordinated gene expression, and regulatory predictability. We identified optimization paths and future directions to accelerate MGE deployment in sustainable agriculture, with possible societal benefits including reduced production costs, increased yield, and improved food and nutritional security.
Side chain engineering in indacenodithiophene- co -benzothiadiazole and its impact on mixed ionic–electronic transport properties
Organic semiconductors are increasingly being decorated with hydrophilic solubilising chains to create materials that can function as mixed ionic–electronic conductors, which are promising candidates for interfacing biological systems with organic electronics. While numerous organic semiconductors, including p- and n-type materials, small molecules and polymers, have been successfully tailored to encompass mixed conduction properties, common to all these systems is that they have been semicrystalline materials. Here, we explore how side chain engineering in the nano-crystalline indacenodithiophene-co-benzothiadiazole (IDTBT) polymer can be used to instil ionic transport properties and how this in turn influences the electronic transport properties. This allows us to ultimately assess the mixed ionic–electronic transport properties of these new IDTBT polymers using the organic electrochemical transistor as the testing platform. Using a complementary experimental and computational approach, we find that polar IDTBT derivatives can be infiltrated by water and solvated ions, they can be electrochemically doped efficiently in aqueous electrolyte with fast doping kinetics, and upon aqueous swelling there is no deterioration of the close interchain contacts that are vital for efficient charge transport in the IDTBT system. Despite these promising attributes, mixed ionic–electronic charge transport properties are surprisingly poor in all the polar IDTBT derivatives. Albeit a “negative” result, this finding clearly contradicts established side chain engineering rules for mixed ionic–electronic conductors, which motivated our continued investigation of this system. We eventually find this anomalous behaviour to be caused by increasing energetic disorder in the polymers with increasing polar side chain content. We have investigated computationally how the polar side chain motifs contribute to this detrimental energetic inhomogeneity and ultimately use the learnings to propose new molecular design criteria for side chains that can facilitate ion transport without impeding electronic transport.
Tunable Cr 4+ Molecular Color Centers
The inherent atomistic precision of synthetic chemistry enables bottom-up structural control over quantum bits, or qubits, for quantum technologies. Tuning paramagnetic molecular qubits that feature optical-spin initialization and readout is a crucial step toward designing bespoke qubits for applications in quantum sensing, networking, and computing. In this work, we demonstrate that the electronic structure that enables optical-spin initialization and readout for S = 1, Cr(aryl) 4 , where aryl = 2,4-dimethylphenyl (1), o-tolyl (2), and 2,3-dimethylphenyl (3), is readily translated into Cr(alkyl) 4 compounds, where alkyl = 2,2,2-triphenylethyl (4), (trimethylsilyl)methyl (5), and cyclohexyl (6). The small ground state zero field splitting values (<5 GHz) for 1–6 allowed for coherent spin manipulation at X-band microwave frequency, enabling temperature-, concentration-, and orientation-dependent investigations of the spin dynamics. Electronic absorption and emission spectroscopy confirmed the desired electronic structures for 4–6, which exhibit photoluminescence from 897 to 923 nm, while theoretical calculations elucidated the varied bonding interactions of the aryl and alkyl Cr 4+ compounds. The combined experimental and theoretical comparison of Cr(aryl) 4 and Cr(alkyl) 4 systems illustrates the impact of the ligand field on both the ground state spin structure and excited state manifold, laying the groundwork for the design of structurally precise optically addressable molecular qubits.
From p- to n-Type Mixed Conduction in Isoindigo-Based Polymers through Molecular Design
Organic mixed ionic and electronic conductors are of significant interest for bioelectronic applications. Here, we use three different isoindigoid building blocks to obtain polymeric mixed conductors with vastly different structural and electronic properties which can be further fine-tuned through the choice of comonomer unit. We show how careful design of the isoindigoid scaffold can afford highly planar polymer structures with high degrees of electronic delocalization, while subtle structural modifications can control the dominant charge carrier (hole or electron) when probed in organic electrochemical transistors. We employ a combination of experimental and computational techniques to probe electrochemical, structural and mixed ionic and electronic properties of the polymer series which in turn allows us to derive important structure-property relations for this promising class of materials in the context of organic bioelectronics. Ultimately, we use these findings to outline robust molecular design strategies for isoindigo-based mixed conductors that can support efficient p-type, n-type and ambipolar transistor operation in an aqueous environment.
Enabling rapid COVID-19 small molecule drug design through scalable deep learning of generative models
We improved the quality and reduced the time to produce machine learned models for use in small molecule antiviral design. Our globally asynchronous multi-level parallel training approach strong scales to all of Sierra with up to 97.7% efficiency. We trained a novel, character-based Wasserstein autoencoder that produces a higher quality model trained on 1.613 billion compounds in 23 minutes while the previous state of the art takes a day on 1 million compounds. Reducing training time from a day to minutes shifts the model creation bottleneck from computer job turnaround time to human innovation time. Our implementation achieves 318 PFLOPs for 17.1% of half-precision peak. We will incorporate this model into our molecular design loop enabling the generation of more diverse compounds; searching for novel, candidate antiviral drugs improves and reduces the time to synthesize compounds to be tested in the lab.
Pandemic drugs at pandemic speed: infrastructure for accelerating COVID-19 drug discovery with hybrid machine learning- and physics-based simulations on high-performance computers
The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottleneck screening the vast number of potential small molecules to shortlist lead compounds for antiviral drug development. New opportunities to accelerate drug discovery lie at the interface between machine learning methods, in this case, developed for linear accelerators, and physics-based methods. The two in silico methods, each have their own advantages and limitations which, interestingly, complement each other. Here, we present an innovative infrastructural development that combines both approaches to accelerate drug discovery. The scale of the potential resulting workflow is such that it is dependent on supercomputing to achieve extremely high throughput. We have demonstrated the viability of this workflow for the study of inhibitors for four COVID-19 target proteins and our ability to perform the required large-scale calculations to identify lead antiviral compounds through repurposing on a variety of supercomputers.
Simulating strongly correlated molecules with a superconducting quantum processor
Many of the biggest challenges in expanding the nation’s access to clean and low-cost energy resources are fundamentally chemistry or materials challenges. An important case is the development of new catalysts for the up-conversion of cheap and readily available materials such as methane or water into materials suitable for use as a fuel such as methanol or oxygen. To understand and exploit such processes, computer simulations of chemical reactions provide a natural complement to experimental studies. Unfortunately, most catalytic reactions involve so-called “strongly correlated” molecules which are notoriously difficult to study with simulation algorithms that can be executed on existing (classical) computers. The recent growth in quantum information science offers an alternative potential route for simulating these difficult systems. As a result, an increasing number of computational chemists are becoming interested in quantum computing. At the same time, quantum information scientists have identified chemistry simulation as a possible first demonstration of a quantum computer providing an improvement over a classical computer. The objective of this project is to accurately simulate strongly correlated molecules on a quantum processor. To meet the high challenges of this objective, new hybrid quantum/classical algorithms will be co-designed with advanced quantum gate developments and computed on customized quantum hardware. Some of the developed techniques will be transferable to study other molecular systems, while the project as a whole will help define better strategies for advancing the quantum simulation of matter more generally.
A Fast Algorithm for Massively Parallel, Long-Term, Simulation of Complex Molecular Dynamics Systems
The advances in theory and computing technology over the last decade have led to enormous progress in applying atomistic molecular dynamics (MD) methods to the characterization, prediction, and design of chemical, biological, and material systems,.
Mechanistic studies of small molecule ligands selective to RNA single G bulges
Abstract Small-molecule RNA binders have emerged as an important pharmacological modality. A profound understanding of the ligand selectivity, binding mode, and influential factors governing ligand engagement with RNA targets is the foundation for rational ligand design. Here, we report a novel class of coumarin derivatives exhibiting selective binding affinity towards single G RNA bulges. Harnessing the computational power of all-atom Gaussian accelerated molecular dynamics simulations, we unveiled a rare minor groove binding mode of the ligand with a key interaction between the coumarin moiety and the G bulge. This predicted binding mode is consistent with results obtained from structure-activity relationship studies and transverse relaxation measurements by nuclear magnetic resonance spectroscopy. We further generated 444 molecular descriptors from 69 coumarin derivatives and identified key contributors to the binding events, such as charge state and planarity, by lasso (least absolute shrinkage and selection operator) regression. Our work deepened the understanding of RNA-small molecule interactions and integrated a new framework for the rational design of selective small-molecule RNA binders.
De novo designed ice-binding proteins from twist-constrained helices
Attaining molecular-level control over solidification processes is a crucial aspect of materials science. To control ice formation, organisms have evolved bewildering arrays of ice-binding proteins (IBPs), but these have poorly understood structure–activity relationships. We propose that reverse engineering using de novo computational protein design can shed light on structure–activity relationships of IBPs. We hypothesized that the model alpha-helical winter flounder antifreeze protein uses an unusual undertwisting of its alpha-helix to align its putative ice-binding threonine residues in exactly the same direction. We test this hypothesis by designing a series of straight three-helix bundles with an ice-binding helix projecting threonines and two supporting helices constraining the twist of the ice-binding helix. Our findings show that ice-recrystallization inhibition by the designed proteins increases with the degree of designed undertwisting, thus validating our hypothesis, and opening up avenues for the computational design of IBPs.
Screening green solvents for multilayer plastic film recycling processes
Multilayer (ML) plastic films are essential packaging materials that help protect products from diverse external factors; however, only 5% of all ML films are recycled in the United States. Solvent-based technologies are a promising alternative for recycling ML films because they enable recovery of constituent polymer resins. For example, the Solvent Targeted Recovery and Precipitation (STRAPTM) process sequentially dissolves and separates polymer components using a series of targeted solvent washes. A crucial design aspect of this process is the impact of selected solvents on human health and on the environment. Here, this work introduces a computational framework that integrates molecular modeling, process modeling, techno-economic analysis (TEA), and life-cycle analysis (LCA) to quickly screen green solvents for solvent-based ML recycling processes. Initial screening for solvents based on selectivity is performed by estimating temperature-dependent solubilities using molecular-scale models. Subsequent screening uses basic estimates of energy use and octanol-water partition coefficients (logP) as key measures of health, safety, and environmental hazards. Detailed process modeling, TEA, and LCA are used on a reduced set of promising solvents identified in early screening steps to more accurately determine how solvent selection and associated operating conditions impact overall economics and environmental impacts. The framework is used for the identification of green solvents (from a database of 1,000 solvents) that separate an industrial ML film composed of polyethylene (PE), ethylene vinyl alcohol (EVOH), and polyethylene terephthalate (PET). Our analysis shows the effectiveness of the framework and reveals fundamental trade-offs between solvent greenness, solubility, and economics. Our work emphasizes the importance of taking a holistic systems view during solvent design and aims to inform the development of new processes for ML film recycling and the identification of new ML films that are easier to recycle.