Machine learning accelerated non-adiabatic molecular dynamics for exciton polaritons
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Predicting the glass transition temperature (T g ) is of critical importance as it governs the thermomechanical performance of conjugated polymers (CPs). Here, we report a predictive modeling framework to predict T g of CPs through the integration of machine learning (ML), molecular dynamics (MD) simulations, and experiments. With 154 T g data collected, an ML model is developed by taking simplified “geometry” of six chemical building blocks as molecular features, where side-chain fraction, isolated rings, fused rings, and bridged rings features are identified as the dominant ones for T g . MD simulations further unravel the fundamental roles of those chemical building blocks in dynamical heterogeneity and local mobility of CPs at a molecular level. The developed ML model is demonstrated for its capability of predicting T g of several new high-performance solar cell materials to a good approximation. The established predictive framework facilitates the design and prediction of T g of complex CPs, paving the way for addressing device stability issues that have hampered the field from developing stable organic electronics.
Future lithium batteries are expected to use solid electrolytes to achieve higher energy density and fast charge capabilities. However, most solid electrolytes are thermodynamically unstable against layered oxide cathodes. In this study, the stability of LiCoO2 (LCO) cathode with Li10GeP2S12 (LGPS) solid electrolyte is investigated using ab initio molecular dynamics (AIMD) and machine learning molecular dynamics (MLMD). The propensity of ionic interdiffusion, formation of a passivating interphase layer, and corresponding decay in cell performance is addressed using a continuum model. Large-scale MLMD simulations confirm that the LCO|LGPS interface permits interdiffusion of cobalt (Co) and other ionic species, leading to the formation and growth of a resistive interphase and to dramatic capacity fade even in the first cycle. We examine the literature evidence that incorporating a thin layer of LiNb0.5Ta0.5O3 (LNTO) between LCO and LGPS prevents the interdiffusion of ions. Atomistic simulations suggest that substituting lithium (Li) in LNTO with Co is thermodynamically unfavorable, thereby inhibiting ionic interdiffusion. The stable Nb5+/Ta5+ states form a rigid metal-oxide framework, which consequently also prevents the substitution of niobium (Nb) or tantalum (Ta). However, continuum-level analysis suggests that the higher mechanical stiffness of LNTO can lead to interfacial delamination between the LCO and LNTO. This phenomenon reduces the effectiveness of the protective layer. This paper, therefore, highlights the need to develop novel interlayers that balance low ionic interdiffusion with low mechanical stiffness.
Mass spectrometry imaging (MSI) plays a pivotal role in investigating the chemical nature of complex systems that underly our understanding in biology and medicine. Multiple fields of life science such as proteomics, lipidomics and metabolomics benefit from the ability to simultaneously identify molecules and pinpoint their distribution across a sample. However, achieving the necessary submicron spatial resolution to distinguish chemical differences between individual cells and generating intact molecular spectra is still a challenge with any single imaging approach. Here, we developed an approach that combines two MSI techniques, matrix-assisted laser desorption/ionization (MALDI) and time-of-flight secondary ion mass spectrometry (ToF-SIMS), one with low spatial resolution but intact molecular spectra and the other with nanometer spatial resolution but fragmented molecular signatures, to predict molecular MSI spectra with submicron spatial resolution. The known relationships between the two MSI channels of information are enforced via a physically constrained machine-learning approach and directly incorporated in the data processing. We demonstrate the robustness of this method by generating intact molecular MALDI-type spectra and chemical maps at ToF-SIMS resolution when imaging mouse brain thin tissue sections. This approach can be readily adopted for other types of bioimaging where physical relationships between methods have to be considered to boost the confidence in the reconstruction product.
Machine learning (ML)-based molecular dynamics (MD) simulations of the formation of a class of N-doped nanoporous carbons are performed to assess their disordered partially graphitized nanoscale structure. The study is motivated by the effectiveness of so-called nitrogen assembly carbons (NACs) for catalysis applications. Benchmark simulations for pure-C disordered graphitic systems reveal the importance of reliably capturing the vdW component of the potentials in order to accurately describe the tendency for layering of disordered graphene-like sheets. In our modeling, this is achieved by a transfer learning strategy incorporating features of the energetics from the optB88-vdW DFT functional into potentials initially trained with a less expensive functional, thereby providing a superior description of the pure-C systems. Generation from MD simulations of realistic partially graphitized structures is significantly more challenging for N-doped versus for pure C systems. However, such structures are achieved by a tailored MD simulation protocol mimicking the experimental synthesis process and in particular incorporating an annealing and subsequent quenching stages. Simulated PXRD patterns effectively reproduce the features of experimental observations for NACs, including the appearance of a prominent but broad (002) peak at around 25, and the development of another weaker feature associated with in-layer ordering of mixed C-N graphene-like sheets.
Simulating molecules and atomic systems at quantum accuracy is a grand challenge for science in the 21 st century. Quantum-accurate simulations would enable the design of new medicines and the discovery of new materials. The defining problem in this challenge is that quantum calculations on large molecules, like proteins or DNA, are fundamentally impossible with current algorithms. In this work, we explore a range of different methods that aim to make large, quantum-accurate simulations possible. We show that using advanced classical models, we can accurately simulate ion channels, an important biomolecular system. We show how advanced classical models can be implemented in an exascale-ready software package. Lastly, we show how machine learning can learn the laws of quantum mechanics from data and enable quantum electronic structure calculations on thousands of atoms, a feat that is impossible for current algorithms. Altogether, this work shows that combining advances in physics models, computing, and machine learning, we are moving closer to the reality of accurately simulating our molecular world.
Ab initio molecular dynamics (AIMD) simulations have become an important tool used in the construction of equations of state (EOS) tables for warm dense matter. Due to computational costs, only a limited number of system state conditions can be simulated, and the remaining EOS surface must be interpolated for use in radiation-hydrodynamic simulations of experiments. In this work, we develop a thermodynamically consistent EOS model that utilizes a physics-informed machine learning approach to implicitly learn the underlying Helmholtz free-energy from AIMD generated energies and pressures. The model, referred to as PIML-EOS, was trained and tested on warm dense polystyrene producing a fit within a 1% relative error for both energy and pressure and is shown to satisfy both the Maxwell and Gibbs–Duhem relations. In addition, we provide a path toward obtaining thermodynamic quantities, such as the total entropy and chemical potential (containing both ionic and electronic contributions), which are not available from current AIMD simulations.
Metal organic framework (MOF)-based mixedmatrix membranes (MMMs), which embed MOF particles in polymer matrices, combine the advantages of polymeric and inorganic membranes. Multiple previous studies have used the Maxwell model together with molecular simulations and machine learning (ML) to predict the performance of MOF/polymer MMMs. However, the assumption of rigid MOF frameworks in molecular simulations limited the accuracy of the data used in the predictions, particularly in predicting molecular diffusivities. We developed a novel workflow integrating ML models with consideration of MOF flexibility to predict the permeability and selectivity of 131,722 MMMs for CO 2 /CH 4 , O 2 /N 2 and He/H 2 separations. The full range of achievable MMM performance within the Maxwell model was analyzed, and several promising MOFs were identified using this workflow. This approach offers an efficient tool for screening any polymer and MOF combination in gas separation applications.
Conotoxins are small and highly potent neurotoxic peptides derived from the venom of marine cone snails which have captured the interest of the scientific community due to their pharmacological potential. These toxins display significant sequence and structure diversity, which results in a wide range of specificities for several different ion channels and receptors. Despite the recognized importance of these compounds, our ability to determine their binding targets and toxicities remains a significant challenge. Predicting the target receptors of conotoxins, based solely on their amino acid sequence, remains a challenge due to the intricate relationships between structure, function, target specificity, and the significant conformational heterogeneity observed in conotoxins with the same primary sequence. We have previously demonstrated that the inclusion of post-translational modifications, collisional cross sections values, and other structural features, when added to the standard primary sequence features, improves the prediction accuracy of conotoxins against non-toxic and other toxic peptides across varied datasets and several different commonly used machine learning classifiers. Here, we present the effects of these features on conotoxin class and molecular target predictions, in particular, predicting conotoxins that bind to nicotinic acetylcholine receptors (nAChRs). We also demonstrate the use of the Synthetic Minority Oversampling Technique (SMOTE)-Tomek in balancing the datasets while simultaneously making the different classes more distinct by reducing the number of ambiguous samples which nearly overlap between the classes. In predicting the alpha, mu, and omega conotoxin classes, the SMOTE-Tomek PCA PLR model, using the combination of the SS and P feature sets establishes the best performance with an overall accuracy (OA) of 95.95%, with an average accuracy (AA) of 93.04%, and an f1 score of 0.959. Using this model, we obtained sensitivities of 98.98%, 89.66%, and 90.48% when predicting alpha, mu, and omega conotoxin classes, respectively. Similarly, in predicting conotoxins that bind to nAChRs, the SMOTE-Tomek PCA SVM model, which used the collisional cross sections (CCSs) and the P feature sets, demonstrated the highest performance with 91.3% OA, 91.32% AA, and an f1 score of 0.9131. The sensitivity when predicting conotoxins that bind to nAChRs is 91.46% with a 91.18% sensitivity when predicting conotoxins that do not bind to nAChRs.
Crystal nucleation is relevant across the domains of fundamental and applied sciences. However, in many cases, its mechanism remains unclear due to a lack of temporal or spatial resolution. To gain insights into the molecular details of nucleation, some form of molecular dynamics simulations is typically performed; these simulations, in turn, are limited by their ability to run long enough to sample the nucleation event thoroughly. To overcome the timescale limits in typical molecular dynamics simulations in a manner free of prior human bias, here, we employ the machine learning-augmented molecular dynamics framework “reweighted autoencoded variational Bayes for enhanced sampling (RAVE).” We study two molecular systems—urea and glycine—in explicit all-atom water, due to their enrichment in polymorphic structures and common utility in commercial applications. From our simulations, we observe multiple back-and-forth nucleation events of different polymorphs from homogeneous solution; from these trajectories, we calculate the relative ranking of finite-sized polymorph crystals embedded in solution, in terms of the free-energy difference between the finite-sized crystal polymorph and the original solution state. We further observe that the obtained reaction coordinates and transitions are highly nonclassical.
Multiscale modeling has a long history of use in structural biology, as computational biologists strive to overcome the time- and length-scale limits of atomistic molecular dynamics. Contemporary machine learning techniques, such as deep learning, have promoted advances in virtually every field of science and engineering and are revitalizing the traditional notions of multiscale modeling. Deep learning has found success in various approaches for distilling information from fine-scale models, such as building surrogate models and guiding the development of coarse-grained potentials. However, perhaps its most powerful use in multiscale modeling is in defining latent spaces that enable efficient exploration of conformational space. In conclusion, this confluence of machine learning and multiscale simulation with modern high-performance computing promises a new era of discovery and innovation in structural biology.
Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.
Here, the LiTaCl 6 solid electrolyte has the lowest activation energy of ionic conduction at ambient conditions (0.165 eV), with a record high ionic conductivity for a ternary compound (11 mS cm –1 ). However, the mechanism has been unclear. We train machine-learning force fields (MLFF) on ab initio molecular dynamics (AIMD) data on-the-fly and perform MLFF MD simulations of AIMD quality up to the nanosecond scale at the experimental temperatures, which allows us to predict accurate activation energy for Li-ion diffusion (at 0.164 eV). Detailed analyses of trajectories and vibrational density of states show that the large-amplitude vibrations of Cl – ions in TaCl 6 – enable the fast Li-ion transport by allowing dynamic breaking and reforming of Li–Cl bonds across the space in between the TaCl 6 – octahedra. We term this process the dynamic-monkey-bar mechanism of superionic Li + transport which could aid the development of new solid electrolytes for all-solid-state lithium batteries.
Not Available
A rapid response is necessary to contain emergent biological outbreaks before they can become pandemics. The novel coronavirus (SARS-CoV-2) that causes COVID-19 was first reported in December of 2019 in Wuhan, China and reached most corners of the globe in less than two months. In just over a year since the initial infections, COVID-19 infected almost 100 million people worldwide. Although similar to SARS-CoV and MERS-CoV, SARS-CoV-2 has resisted treatments that are effective against other coronaviruses. Crystal structures of two SARS-CoV-2 proteins, spike protein and main protease, have been reported and can serve as targets for studies in neutralizing this threat. We have employed molecular docking, molecular dynamics simulations, and machine learning to identify from a library of 26 million molecules possible candidate compounds that may attenuate or neutralize the effects of this virus. The viability of selected candidate compounds against SARS-CoV-2 was determined experimentally by biolayer interferometry and FRET-based activity protein assays along with virus-based assays. In the pseudovirus assay, imatinib and lapatinib had IC 50 values below 10 μM, while candesartan cilexetil had an IC 50 value of approximately 67 µM against M pro in a FRET-based activity assay. Comparatively, candesartan cilexetil had the highest selectivity index of all compounds tested as its half-maximal cytotoxicity concentration 50 (CC 50 ) value was the only one greater than the limit of the assay (>100 μM).
Zeolites are the main solid catalysts used by the chemical industry. The use of zeolites in separations and as shape selective catalysts requires control of the width and connectivity of their pores. 235 distinct zeolite frameworks have been synthesized to date, of over 2 million that have been proposed. Recent work indicates that the limitation is in large part kinetic: new synthetic pathways are required to access new zeolites. Organic cations are used to direct the synthesis towards specific zeolites. However, the molecular mechanisms by which cations direct the nucleation towards specific zeolites is not known. Elucidating these mechanisms is key to realize new zeolites for catalysis and separations, and is the focus of this project. This project developed and implemented a synergistic, data-driven computational and experimental approach to resolve the molecular pathways of nucleation, growth, and polymorph selection of zeolites and the role of organic cations in directing their formation. The project developed computationally efficient and accurate models for the study of the nucleation and growth of pure silica zeolites in molecular simulations, using machine learning with data from experiments. Simulations with these models were integrated with scanning tunneling electron microscopy, computer vision, and deep learning to unveil the molecular pathways of formation of a zeolite. Of particular interest in this project was to elucidate the role of amorphous precursors in the nucleation of the zeolite. Previous experiments indicate that zeolites are born within non-crystalline aggregates in which the silicates and organic cations have local and medium range order similar to that of the zeolite. The organic cations that direct the formation of zeolites and those that direct the formation of ordered mesoporous silicas are similar. We hypothesized that the frustrated attraction that for large organic cations leads to the formation of stable mesophases that direct the synthesis of mesoporous silicas, could promote the formation of metastable mesophases that can assist in the nucleation and polymorph selection of zeolites. The simulations resolved how structure directing agents build crystalline order and showed that mesoscopic pre-ordering occurs is not required to facilitate the nucleation of zeolites, because the synthesis occurs at high driving forces, where the barriers for nucleation are negligible. This project unveiled that polymorph selection in zeolite synthesis occurs after nucleation, opening a distinct area of control through the kinetics of growth and not through nucleation barriers.
Molecular simulations have provided valuable insight into the microscopic mechanisms underlying homogeneous ice nucleation. While empirical models have been used extensively to study this phenomenon, simulations based on first-principles calculations have so far proven prohibitively expensive. Here, we circumvent this difficulty by using an efficient machine-learning model trained on density-functional theory energies and forces. We compute nucleation rates at atmospheric pressure, over a broad range of supercoolings, using the seeding technique and systems of up to hundreds of thousands of atoms simulated with ab initio accuracy. The key quantity provided by the seeding technique is the size of the critical cluster (i.e., a size such that the cluster has equal probabilities of growing or melting at the given supersaturation), which is used together with the equations of classical nucleation theory to compute nucleation rates. We find that nucleation rates for our model at moderate supercoolings are in good agreement with experimental measurements within the error of our calculation. We also study the impact of properties such as the thermodynamic driving force, interfacial free energy, and stacking disorder on the calculated rates.