Engineering PapersSearch

SEARCH · Engineering Papers

Results for “MLIPs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Application-specific machine-learned interatomic potentials: exploring the trade-off between DFT convergence, MLIP expressivity, and computational cost

Machine-learned interatomic potentials (MLIPs) are revolutionizing computational materials science and chemistry by offering an efficient alternative to ab initio molecular dynamics (MD) simulations. However, fitting high-quality MLIPs remains a challenging, time-consuming, and computationally intensive task where numerous trade-offs have to be considered, e.g., How much and what kind of atomic configurations should be included in the training set? Which level of ab initio convergence should be used to generate the training set? Which loss function should be used for fitting the MLIP? Which machine learning architecture should be used to train the MLIP? The answers to these questions significantly impact both the computational cost of MLIP training and the accuracy and computational cost of subsequent MLIP MD simulations. In this study, we use a configurationally diverse beryllium dataset and quadratic spectral neighbor analysis potential. We demonstrate that joint optimization of energy versus force weights, training set selection strategies, and convergence settings of the ab initio reference simulations, as well as model complexity can lead to a significant reduction in the overall computational cost associated with training and evaluating MLIPs. This opens the door to computationally efficient generation of high-quality MLIPs for a range of applications which demand different accuracy versus training and evaluation cost trade-offs.

36 MATERIALS SCIENCE

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

Multi-fidelity learning for interatomic potentials: low-level forces and high-level energies are all you need

The promise of machine learning interatomic potentials (MLIPs) has led to an abundance of public quantum mechanical (QM) training datasets. The quality of an MLIP is directly limited by the accuracy of the energies and atomic forces in the training dataset. Unfortunately, most of these datasets are computed with relatively low-accuracy QM methods, e.g. density functional theory with a moderate basis set. Due to the increased computational cost of more accurate QM methods, e.g. coupled-cluster theory with a complete basis set (CBS) extrapolation, most high-accuracy datasets are much smaller and often do not contain atomic forces. The lack of high-accuracy atomic forces is quite troubling, as training with force data greatly improves the stability and quality of the MLIP compared to training to energy alone. Because most datasets are computed with a unique level of theory, traditional single-fidelity (SF) learning is not capable of leveraging the vast amounts of published QM data. In this study, we apply multi-fidelity learning (MFL) to train an MLIP to multiple QM datasets of different levels of accuracy, i.e. levels of fidelity. Specifically, we perform three test cases to demonstrate that MFL with both low-level forces and high-level energies yields an extremely accurate MLIP—far more accurate than a SF MLIP trained solely to high-level energies and almost as accurate as a SF MLIP trained directly to high-level energies and forces. Therefore, MFL greatly alleviates the need for generating large and expensive datasets containing high-accuracy atomic forces and allows for more effective training to existing high-accuracy energy-only datasets. Indeed, low-accuracy atomic forces and high-accuracy energies are all that are needed to achieve a high-accuracy MLIP with MFL.

36 MATERIALS SCIENCE

Application of machine learning interatomic potentials in heterogeneous catalysis

Heterogeneous catalysts are crucial in modern societies as they promote sustainability by enabling lower-energy pathways for various chemical reactions. While Density Functional Theory (DFT) computations can provide critical insights into how heterogeneous catalysts operate at the atomic level, they are limited by computational costs and unfavorable scaling with system size. Recently, machine learning interatomic potentials (MLIPs) have emerged as a promising alternative to DFT, offering near-DFT accuracy at significantly reduced cost. Here, in this perspective, we discuss the application of MLIPs in heterogeneous catalyst modeling as a surrogate for DFT. We detail how MLIPs have been applied in thermal catalysis to probe active sites, enable studying complex metallic and nanoporous catalysts, and investigate the reconstruction of catalytic surfaces. We review the use of MLIPs in electrocatalysis and photocatalysis, emphasizing their capabilities in studying transition metal oxide surfaces and solid–liquid interfaces. We also discuss the current limitations of MLIPs, particularly their challenges with transferability and description of non-local interactions. Finally, we conclude by identifying promising and underexplored domains in which MLIPs can further advance our understanding of heterogeneous catalysts.

Catalytic surfaces

Python Library for Monte Carlo Simulations with Ab Initio and Machine-Learned Interatomic Potentials

There is a growing need in the simulation community for software that provides a transparent, reproducible, usable, and extensible (TRUE) Monte Carlo (MC) simulation framework employing energies from ab initio methods and machine-learning interatomic potentials (MLIPs). We introduce a Python library (ASE-MC) that adds Monte Carlo functionality to the Atomic Simulation Environment (ASE) package. Now, we can combine the powerful tools used to build systems and perform ab initio and MLIP in ASE with MC simulation algorithms to sample the configurational space with a concise Python script. After presenting the design philosophy, we demonstrate the flexibility of our approach using selected examples. These example simulations include liquid water described with a message-passing MLIP in the canonical and isothermal–isobaric ensembles, sampling the characteristic dihedral angle of biphenyl and comparing an MLIP to first-principles calculations, and a grand canonical Monte Carlo simulation of ammonia adsorption on Pt(111). These examples showcase the main features of the software, which include flexibility in the choice of ab initio or MLIP engine, ab initio or MLIP grand canonical MC with cavity bias insertions and deletions, the ability to add custom MC moves to the move set, and how users can condense complex MC workflows into a single Python script. Finally, this library serves as a framework for reproducible Monte Carlo simulations, facilitating easy reproduction of the work and application to new systems.

97 MATHEMATICS AND COMPUTING

Full-stack Quantification of Variability in Predicting Ion Transport Properties using Machine-learned Interatomic Potentials

Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is therefore crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods, and improving the MD sampling statistics.

36 MATERIALS SCIENCE

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics

Prediction and Experimental Verification of Electrolyte Solvation Structure from an OMol25-Trained Interatomic Potential

A molecular-level understanding of electrolyte solvation structure and ion–ion correlations is critical to developing next-generation battery chemistries. Atomistic simulation capabilities with sufficient accuracy, speed, and transferability to deliver reliable structural insights while avoiding arduous system-specific reparameterization are thus highly desirable. Machine learning interatomic potentials (MLIPs) trained on large, chemically diverse data sets are revolutionizing computational chemistry, enabling molecular dynamics simulations of battery electrolytes with near-DFT accuracy over 10,000× faster than DFT. While previous MLIP training data sets with suitable elemental coverage for electrolytes have been based on inorganic materials, the Open Molecules 2025 (OMol25) data set provides large-scale molecular DFT MLIP training data with broad elemental coverage and specifically samples tens of millions of electrolyte configurations. Here, we integrate computational modeling with experimental validation to systematically assess the ability of large-scale MLIPs pretrained on materials data or on OMol25 to accurately resolve nanoscale structural organization and ion-solvation characteristics in Na-ion battery electrolytes across diverse physicochemical conditions and compositional regimes. We find that the OMol25-trained Universal Model of Atoms (UMA-OMol) predicts experimentally measured densities and X-ray structure factors in substantially better agreement compared to state-of-the-art models trained only on inorganic materials data. Using UMA-OMol, we further analyze systematic trends in solvation structure as a function of cation identity, anion chemistry, salt concentration, and solvent topology. We observe that increasing system temperature amplifies the heterogeneity within the solvation environment, perturbing cation–solvent interactions and promoting the formation of contact ion pairs (CIPs). Moreover, subtle variations in the solvent topology of glyme-based electrolytes cause pronounced changes in ion correlations and solvation structure. The experimental agreement and microscopic insights shown here position OMol25-trained MLIPs as a practical route to predictive, high-throughput electrolyte simulations beyond the limits of classical force fields and direct DFT molecular dynamics, serving as a powerful tool for accelerating the design of next-generation Na-ion battery electrolytes and beyond.

MLIPs

Improving Bond Dissociations of Reactive Machine Learning Potentials through Physics-Constrained Data Augmentation

In the field of computational chemistry, predicting bond dissociation energies (BDEs) presents well-known challenges, particularly due to the multireference character of reactive systems. Many chemical reactions involve configurations where single-reference methods fall short, as the electronic structure can significantly change during bond breaking. As generating training data for partially broken bonds is a challenging task, even state-of-the-art reactive machine learning interatomic potentials (MLIPs) often fail to predict reliable BDEs and smooth dissociation curves. By contrast, simple and inexpensive physics-based models, such as the well-established Morse potential, do not suffer from any such limitations. This work leverages the Morse potential to improve reactive MLIPs by augmenting the training data set with inexpensive Morse data along the dissociation pathways. Further, this physics-constrained data augmentation (PCDA) approach results in MLIPs with smooth bond dissociation curves as well as near coupled-cluster level BDEs, all without requiring any expensive multireference quantum mechanical calculations. A case study for methane combustion demonstrates how the PCDA approach can improve an existing reactive MLIP, namely, ANI-1xnr. In conclusion, not only are the BDEs and bond dissociation curves for all radicals and molecules significantly improved compared to ANI-1xnr but the PCDA-trained MLIP retains the reliability of ANI-1xnr when performing reactive molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Development of Machine-Learned Interatomic Potentials to Predict Structure, Transport, and Reactivity in Platinum-Based Fuel Cells

Machine-learned interatomic potentials (MLIPs) have rapidly progressed in accuracy, speed, and data efficiency in recent years. However, training robust MLIPs in multicomponent systems remains a challenge. In this work, we train an MLIP to describe hydrated Nafion ionomers and platinum catalysts, which are important components of fuel cells, by constructing a diverse training set to describe the bulk polymer and interfacial catalyst–polymer interactions well. We use our trained MLIP to study the properties of the platinum–Nafion system, including polymer structure, proton mobility in a bulk Nafion polymer and near a platinum-Nafion interface, and reactions near and far from the interface, finding excellent results for structure and reactions contained within our training set. Transport seems to be well described, with both vehicular transport and Grotthuss hopping captured, although converged calculations of diffusivities were not computed because they require calculations of tens of nanoseconds that are challenging with current state-of-the-art MLIPs. The combined insights that this model provides can be leveraged to optimize fuel cell performance, and the approach can be applied to other chemical processes and devices where structure, transport, and reactivity all contribute to the overall observed performance.

33 ADVANCED PROPULSION SYSTEMS

Teacher-student training improves the accuracy and efficiency of machine learning interatomic potentials

Machine learning interatomic potentials (MLIPs) are revolutionizing the field of molecular dynamics (MD) simulations. Recent MLIPs have tended towards more complex architectures trained on larger datasets. The resulting increase in computational and memory costs may prohibit the application of these MLIPs to perform large-scale MD simulations. Herein, we present a teacher-student training framework in which the latent knowledge from the teacher (atomic energies) is used to augment the students' training. We show that the light-weight student MLIPs have faster MD speeds at a fraction of the memory footprint compared to the teacher models. Remarkably, the student models can even surpass the accuracy of the teachers, even though both are trained on the same quantum chemistry dataset. Our work highlights a practical method for MLIPs to reduce the resources required for large-scale MD simulations.

36 MATERIALS SCIENCE

Modeling the behavior of concentrated aqueous HNO 3 using machine learning interatomic potentials

We develop two multi-defect machine learning interatomic potentials (MLIPs) trained at the BLYP-D2 and PBE-D3 density functional theories using the DeepMD-kit, allowing for the investigation of structural and thermodynamic properties of nitric acid over a wide range of concentrations via molecular dynamics (MD) simulations. We directly compute the degree of dissociation, α, and pK a from MD simulations, revealing that HNO 3 behaves as a weaker acid at higher concentrations, noting that our standard-state pK a value is in excellent agreement with the experimental one. In general, good agreement is observed with experimental results such as α and density outside the training dataset, with only modest deviations at low-to-medium concentrations. We benchmark our custom multi-defect DeepMD MLIPs against foundational models MACE-MP0 and MACE-OFF23. The foundation models capture some aspects of HNO 3 /NO 3 − solvation in concentrated nitric acid but show noticeable density errors and miss subtle structural features relevant to spectroscopy, whereas the bespoke DeepMD MLIPs yield more compact solvation shells, reproduce density-concentration trends, and run ∼12–15× faster than MACE-MP0. Although classical FFs are still more efficient and match experimental densities better, they lack chemical reactivity and thus cannot predict α or pK a , underscoring the need for system-specific reactive MLIPs beyond universal MLIPs.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation

Efficient machine learning interatomic potentials robust for liquid and multiple solid polymorphs of NaF and KF

Achieving atomic-level understanding of crystallization of molten salts is of importance to a wide range of technological applications. Recent work [Fan et al., Proc. Natl. Acad. Sci. USA 122, e2425702122 (2025)] revealed that crystal nucleation in molten LiF salt is a multistage process according to the molecular-dynamics (MD) simulations based on an atomic cluster expansion (ACE) machine-learning interatomic potential (MLIP). In order to understand the influence of increasing cation size on nucleation pathways and nucleation rates of molten fluoride salts, here we develop two new ACE MLIPs for NaF and KF. The two ACE MLIPs feature DFT-SCAN-level accuracy for liquid and multiple solid polymorphs over a wide temperature (0–2000 K) and pressure (0–100 GPa) range, and also reproduce well a number of experimental data for solid and liquid equilibrium properties. The efficiency of the two ACE MLIPs enable million-atom-scale or microsecond-scale MD simulations. The two general-purpose ACE MLIPs are expected to be useful for atomistic simulations for different purposes, in addition to studying crystallization of molten NaF and KF salts.

Crystal melting

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage

Effects of Composition and Oxidation States on the Structures of Chromium-Containing Sodium Silicate Glasses: Molecular Dynamics Simulations using Machine Learning Interatomic Potentials

Chromium represents a significant challenge for the vitrification of high-level nuclear waste into silicate and borosilicate glasses due to its low solubility and variable oxidation states, which can limit the waste loading due to promotion of crystallization or phase separation during processing. In this study, we modeled chromium containing silicate glasses using molecular dynamics simulations with three machine learning interatomic potentials (MLIPs), MACE, CHGNet, and PFP were employed, to gain insights on glass composition and oxidation states on the structures of these glasses. One of the goals is to evaluate their ability of these MLIPs to accurately represent the general structure of silicate glasses and chromium local environments as a function of chromium oxidation states. Density Functional Theory (DFT) based calculations and experimental data such as neutron structure factors were used to validate the structural models. It was found that the foundation models of all three MLIPs are able to reproduce general structural features of the sodium silicate glass structure consistent with experimental and DFT data, but only CHGNet and PFP can accurately capture the oxidation states and local environment of chromium: tetrahedral for Cr6+ and octahedral for Cr3+. Furthermore, we studied the effect of varying Cr3+/ Cr6+ (Cr3+/Crtotal) ratio and total chromium content using PFP. Our results show that Cr6+ enhances network polymerization by reducing non-bridging oxygens through Na? charge compensation required due to the formation of chromate (CrO42-) species, while Cr³? acts as a network modifier that disrupts connectivity. System size effects on the structural characteristics and chromium environments were also tested using the PFP potential. This work highlights the importance of careful validation on the precision, transferability, and potential of MLIPs for modeling glasses containing transition metal elements that can exist in multiple oxidation states. It is also encouraging to see the foundational models are all three MLFFs are able to reproduce the basic sodium silicate glass structures, while suggesting additional training or refining is needed to improve the description of more complex systems containing transition metals.

Puga, Christina L.

A Universal Augmentation Framework for Long-Range Electrostatics in Machine Learning Interatomic Potentials

Most current machine learning interatomic potentials (MLIPs) rely on short-range approximations, without explicit treatment of long-range electrostatics. To address this, we recently developed the Latent Ewald Summation (LES) method, which infers electrostatic interactions, polarization, and Born effective charges (BECs), just by learning from energy and force training data. Here, in this study, we present LES as a standalone library, compatible with any short-range MLIP, and demonstrate its integration with methods such as MACE, NequIP, Allegro, CACE, CHGNet, and UMA. We benchmark LES-enhanced models on distinct systems, including bulk water, polar dipeptides, and gold dimer adsorption on defective substrates, and show that LES not only captures correct electrostatics but also improves accuracy. Additionally, we scale LES to large and chemically diverse data by training MACELES-OFF on the SPICE set containing molecules and clusters, making a universal MLIP with electrostatics for organic systems, including biomolecules. MACELES-OFF is more accurate than its short-range counterpart (MACE-OFF) trained on the same data set, predicts dipoles and BECs reliably, and has better descriptions of bulk liquids. By enabling efficient long-range electrostatics without directly training on electrical properties, LES paves the way for electrostatic foundation MLIPs.

Kim, Dongjin [University of California, Berkeley,

Liquid–Vapor Phase Equilibrium in Molten Aluminum Chloride (AlCl 3 ) Enabled by Machine Learning Interatomic Potentials

Molten salts are promising candidates in numerous clean energy applications, where knowledge of thermophysical properties and vapor pressure across their operating temperature ranges is critical for safe operations. Due to challenges in evaluating these properties using experimental methods, fast and scalable molecular simulations are essential to complement the experimental data. In this study, we developed machine learning interatomic potentials (MLIP) to study the AlCl 3 molten salt across varied thermodynamic conditions (T = 473–613 K and P = 2.7–23.4 bar), which allowed us to predict temperature-surface tension correlations and liquid–vapor phase diagram from direct simulations of two-phase coexistence in this molten salt. Two MLIP architectures, a Kernel-based potential and neural network interatomic potential (NNIP), were considered to benchmark their performance for AlCl 3 molten salt using experimental structure and density values. The NNIP potential employed in two-phase equilibrium simulations yields the critical temperature and critical density of AlCl 3 that are within 10 K (∼3%) and 0.03 g/cm 3 (∼7%) of the reported experimental values. An accurate correlation between temperature and viscosities is obtained as well. In doing so, we report that the inclusion of low-density configurations in their training is critical to more accurately represent the AlCl 3 system across a wide phase-space. The MLIP trained using PBE-D3 functional in the ab initio molecular dynamics (AIMD) simulations (120 atoms) also showed close agreement with experimentally determined molten salt structure comprising Al 2 Cl 6 dimers, as validated using Raman spectra and neutron structure factor. Furthermore, the PBE-D3 as well as its trained MLIP showed better liquid density and temperature correlation for AlCl 3 system when compared to several other density functionals explored in this work. Overall, the demonstrated approach to predict temperature correlations for liquid and vapor densities in this study can be employed to screen nuclear reactors-relevant compositions, helping to mitigate safety concerns.

Ab initio molecular dynamics