Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning force fields”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Quasi-Classical Trajectory Calculation of Rate Constants Using an Ab Initio Trained Machine Learning Model (aML-MD) with Multifidelity Data

Machine learning (ML) provides a great opportunity for the construction of models with improved accuracy in classical molecular dynamics (MD). However, the accuracy of a ML trained model is limited by the quality and quantity of the training data. Generating large sets of accurate ab initio training data can require significant computational resources. Furthermore, inconsistent or incompatible data with different accuracies obtained using different methods may lead to biased or unreliable ML models that do not accurately represent the underlying physics. Recently, transfer learning showed its potential for avoiding these problems as well as for improving the accuracy, efficiency, and generalization of ML models using multifidelity data. In this work, ab initio trained ML-based MD (aML-MD) models are developed through transfer learning using DFT and multireference data from multiple sources with varying accuracy within the Deep Potential MD framework. Further, the accuracy of the force field is demonstrated by calculating rate constants for the H + HO 2 → H 2 + 3 O 2 reaction using quasi-classical trajectories. We show that the aML-MD model with transfer learning can accurately predict the rate constants while reducing the computational cost by more than five times compared to the use of more expensive quantum chemistry training data sets. Hence, the aML-MD model with transfer learning shows great potential in using multifidelity data to reduce the computational cost involved in generating the training set for these potentials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dark QCD: the Next Frontier in Dark Matter

There has been a surge of interest in hidden valley models with new, strong forces, sometimes called "dark QCD". These models propose asymmetric, composite dark matter in the form of "dark hadrons" that would evade direct and indirect bounds as well as typical collider DM searches for large missing transverse momentum accompanied by radiation. However, evidence of these models can still be found in collider datasets by targeting their unique phenomenological signatures, which include semi visible jets, emerging jets, and soft unclustered energy patterns. We will present the first experimental results for all of these signatures, which have made significant strides in exploring the vast space of dark QCD models. We will further discuss the prospects for dramatic improvements in sensitivity using machine learning.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

From bulk to surface: Structure and dynamics of amorphous alumina from deep potential molecular dynamics

Understanding the atomic-scale structure and dynamics of amorphous oxide surfaces is essential for interpreting their chemical reactivity, mechanical stability, and interfacial behavior, yet direct experimental characterization remains challenging. We employ Deep Potential (DP) molecular dynamics to generate large-scale, ab initio -quality models of amorphous Al 2 O 3 bulk glasses and melt-quenched free surfaces, enabling a quantitative analysis of both structure and relaxation dynamics with statistical confidence inaccessible to direct ab initio simulation. The trained DP model reproduces experimental liquid and glass structure, captures the cooling-rate dependence of the bulk glass transition, and corrects systematic biases in the polyhedral populations predicted by widely used classical force fields. At the free surface, mass density recovers to bulk values over ~10 Å, while local coordination requires a slightly wider subsurface region to fully converge. The outermost layer is oxygen-enriched, exhibits altered polyhedral connectivity with contracted Al–O bonds, and hosts a broad population of under-coordinated motifs (notably AlO 3 and OAl 2 ) whose abundances are governed by glass stability. These under-coordinated surface motifs exhibit distinct vibrational signatures and occur as locally paired Lewis acid and Brønsted base sites consistent with bond-valence compensation, yet remain spatially dispersed rather than aggregating into extended clusters. Despite this pronounced structural heterogeneity, surface relaxation and the glass-transition temperature remain comparable to their bulk counterparts, suggesting that the disordered surface is kinetically stable once formed. Together, these results establish a molecular-level picture of amorphous alumina surfaces and demonstrate the capability of machine-learned potentials to resolve structure–property relationships in disordered oxide interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multiscale Modeling Meets Machine Learning: What Can We Learn?

Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.

machine learning, multiscale modeling, physics-bas↗

Parton labeling without matching: unveiling emergent labelling capabilities in regression models

Parton labeling methods are widely used when reconstructing collider events with top quarks or other massive particles. State-of-the-art techniques are based on machine learning and require training data with events that have been matched using simulations with truth information. In nature, there is no unique matching between partons and final state objects due to the properties of the strong force and due to acceptance effects. We propose a new approach to parton labeling that circumvents these challenges by recycling regression models. The final state objects that are most relevant for a regression model to predict the properties of a particular top quark are assigned to said parent particle without having any parton-matched training data. This approach is demonstrated using simulated events with top quarks and outperforms the widely used $χ$ 2 method.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Deep learning interfacial momentum closures in coarse-mesh CFD two-phase flow simulation using validation data

Multiphase flow phenomena have been widely observed in the industrial applications while it remains a challenging yet unsolved problems. Three-dimensional computational fluid dynamics (CFD) approaches resolve the flow fields on a finer special and temporal scales which can complement the dedicated experimental study. However, closures have to be introduced to reflect the underlying physics in multiphase flow. Among them, the interfacial forces, including drag, lift, turbulent dispersion and wall lubrication forces, play in important role on the bubble’s distribution and migration in liquid-vapor two-phase flow. Development of those closures traditionally rely on the experimental data and analytical derivation with simplified assumptions which usually cannot deliver a universal solution across wide range of flow conditions. In this paper, a data-driven approach, named as Feature Similarity Measurement (FSM), is developed and applied to improve the simulation capability of two-phase flow with coarse-mesh CFD approach. Interfacial momentum transfer in adiabatic bubbly flow serves as the focus of the present study. Both a mature and a simplified set of interfacial closures are taken as the low fidelity data. Experimental data and fine mesh CFD simulations results are adopted as high-fidelity data. Qualitative and quantitative analysis are performed in this paper which reveals that FSM can substantially improve the prediction of coarse mesh CFD model regardless of the choice of interfacial closures and it provides scalability and consistency across discontinuous flow regimes. Furthermore, it demonstrates that data-driven method can aid the multiphase flow modeling by exploring the connections between local physical features and simulation errors.

97 MATHEMATICS AND COMPUTING↗

Unusual dynamics of tetrahedral liquids caused by the competition between dynamic heterogeneity and structural heterogeneity

Tetrahedral liquids exhibit intriguing thermodynamic and transport properties because of the various ways tetrahedra can be packed and connected. Recently, an unusual temperature dependence of the stretching exponent β in a model tetrahedral liquid ZnCl 2 from T m + 85 K to T m + 35 K has been reported using neutron-spin echo spectroscopy. This discovery stands in sharp contrast to other glass-forming liquids. In this study, we conducted neural network force field driven molecular dynamic simulations of ZnCl 2 . We found a non-monotonic temperature dependence of β from liquid to supercooled liquid temperatures. Further structural decomposition and dynamic analysis suggest that this unusual dynamic behavior is a result of the competition between the decrease in the diversity of tetrahedra motifs (structural heterogeneity) and the increase in glassy dynamic heterogeneity. Furthermore, this result may contribute to new understandings of the structural relaxation of other network liquids.

36 MATERIALS SCIENCE↗

Physics-Informed and Data-Driven Prediction of Residual Stress in Three-Dimensional Machining

Efficient and reliable prediction of machining-induced residual stress (RS) is a key requirement for truly integrated computational materials engineering (ICME). Currently available process modeling approaches, including empirical, analytical, and numerical methodologies lack predictive power and require substantial calibration and validation data. Moreover, most model-based approaches consider only two-dimensional (2D) (i.e., orthogonal), cutting processes. Meanwhile, industrial processes such as milling, turning, and drilling are inherently three-dimensional (3D). The present work attempts to bridge the gap between 2D and 3D through careful consideration of the process physics, including geometric, kinematic, and size-effect constraints to realize robust prediction of how RS develops in 3D machining. Using a novel in-situ experimental technique and digital image correlation (DIC) to determine equivalent Hertzian contact widths, contact pressures, and friction coefficients, the proposed methodology leverages a discretized conversion algorithm that includes multi-pass shakedown effects. This paper presents a semi-analytical model to predict machining-induced RS in 3D turning operations, which are used representatively for 3D processes more generally. Rather than follow a ‘brute force’ 3D FEM approach or conduct countless experiments to train a purely data-driven machine learning algorithm, the proposed approach builds on previous 2D modeling work. Through careful consideration of the process physics, including complex geometry/kinematic considerations of 3D turning, the authors demonstrated an experimentally calibrated approach, as well as validation based on published RS data. Model predictions and previously published measurement data of RS depth profiles for turning of Inconel 718 were compared for a range of process parameters. Correlation between the proposed 3D model and validation data was found to be within the margin of experimental error for most conditions. The proposed model appears to capture the overall behavior of 3D RS depth profiles with acceptable accuracy, particularly the key metrics of near-surface stress, peak stress magnitude and location, as well as overall stress profile depth. This report presents a physics-informed, data-driven approach for efficient calibration of a 2D model for machining-induced RS through DIC analysis of in-situ characterized subsurface displacement fields.

42 ENGINEERING↗

DER Cybersecurity Detection and Response Suite

SAND2024-08475O The Distributed Energy Resource (DER) Cybersecurity Detection and Response Suite is a solution for distributed energy resource (DER) systems. The DER Security Orchestration, Automation, and Response (SOAR) solution that uses alerts from signature- and behavior-based Intrusion Detection Systems are intended to be deployed as bump-in-the-wire (BITW) devices in front of DER equipment. The fielded application would use multiple intrusion detection systems that report data to SOAR to respond to cyberattacks. The suite consists of two software components: • The proactive intrusion detection and mitigation system (PIDMS) secures grid-edge photovoltaic smart inverters and other equipment in distributed energy resource systems. It is a distributed BITW solution; cyber and physical data are automatically processed using network inspection tools and custom machine learning algorithms to detect abnormal events and correlate cyber-physical events. • The Security Orchestration, Automation, and Response for Distributed Energy Resources (SOAR4DER) application ingests data from several intrusion detection systems to quickly block attacks and revert DER systems to good states. Using a collection of intrusion detection system technologies on a BITW device, it incorporates physical and cyber data to detect abnormal and potential malicious behaviors. Multiple SOAR playbooks then use the intrusion detection system data streams to automatically defend the system. SOAR4DER system testing showed detection and response times under 30 seconds for all adversary reconnaissance, denial-of-service attacks, malicious Modbus commands, brute-force logins, and machine-in-the-middle attacks. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, Jay↗

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION↗

How Well Do We Know the Neutron-Matter Equation of State at the Densities Inside Neutron Stars? A Bayesian Approach with Correlated Uncertainties

Here, we introduce a new framework for quantifying correlated uncertainties of the infinite-matter equation of state derived from chiral effective field theory (𝜒⁢EFT ). Bayesian machine learning via Gaussian processes with physics-based hyperparameters allows us to efficiently quantify and propagate theoretical uncertainties of the equation of state, such as 𝜒⁢EFT truncation errors, to derived quantities. We apply this framework to state-of-the-art many-body perturbation theory calculations with nucleon-nucleon and three-nucleon interactions up to fourth order in the 𝜒⁢EFT expansion. This produces the first statistically robust uncertainty estimates for key quantities of neutron stars. We give results up to twice nuclear saturation density for the energy per particle, pressure, and speed of sound of neutron matter, as well as for the nuclear symmetry energy and its derivative. At nuclear saturation density, the predicted symmetry energy and its slope are consistent with experimental constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

Assessing entropy for catalytic processes at complex reactive interfaces

When chemical reactions are accelerated by a catalyst, entropy differences between reactants and their transient intermediates can be the driving force behind the promotion or inhibition of desired and parasitic chemical pathways. Understanding and controlling catalytic processes therefore requires both a fundamental and practicable understanding of entropy in addition to enthalpy. In unstructured media such as the vapor phase equilibrated with sparsely covered surfaces, entropy can be adequately accounted for by well-established approaches based on translational, rotational, and harmonic vibrational partition functions. However, these approximations become inadequate in more complex condensed phase environments, e.g., solid liquid interfaces of confined reaction spaces. In this chapter, we provide an overview of the state-of-art in the computational quantification of entropy and its known ramifications on catalysis. The fundamental roles of thermodynamics and kinetics in catalysis are covered in enough detail to appreciate and contextualize the computational methods employed to compute chemically accurate estimates of entropy. These methods are discussed in appropriate detail and range from the ubiquitous harmonic oscillator approximation where entropy unrelated to high frequency oscillations is typically underestimated, to enhanced free energy sampling with molecular dynamics where the desired accuracy must be weighed against the associated computational cost of obtaining it. The rising importance of machine learning and artificial intelligence in accelerating methodological progress in this field is touched upon, as well. Finally, applications, successes, and pitfalls of using these methods are provided to showcase past and present accomplishments while clarifying where improvements in both understanding and methodology are still needed.

Kollias, Loukas↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Quantifying uncertainties and correlations in the nuclear-matter equation of state

We perform statistically rigorous uncertainty quantification (UQ) for chiral effective field theory (χ EFT) applied to infinite nuclear matter up to twice nuclear saturation density. The equation of state (EOS) is based on high-order many-body perturbation theory calculations with nucleon-nucleon and three-nucleon interactions up to fourth order in the χ EFT expansion. From these calculations our newly developed Bayesian machine-learning approach extracts the size and smoothness properties of the correlated EFT truncation error. Furthermore, we then propose a novel extension that uses multitask machine learning to reveal correlations between the EOS at different proton fractions. The inferred in-medium χ EFT breakdown scale in pure neutron matter and symmetric nuclear matter is consistent with that from free-space nucleon-nucleon scattering. These significant advances allow us to provide posterior distributions for the nuclear saturation point and propagate theoretical uncertainties to derived quantities: the pressure and incompressibility of symmetric nuclear matter, the nuclear symmetry energy, and its derivative. Our results, which are validated by statistical diagnostics, demonstrate that an understanding of truncation-error correlations between different densities and different observables is crucial for reliable UQ. The methods developed here are publicly available as annotated Jupyter notebooks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The suitability of differentiable, physics-informed machine learning hydrologic models for ungauged regions and climate change impact assessment

As a genre of physics-informed machine learning, differentiable process-based hydrologic models (abbreviated as δ or delta models) with regionalized deep-network-based parameterization pipelines were recently shown to provide daily streamflow prediction performance closely approaching that of state-of-the-art long short-term memory (LSTM) deep networks. Meanwhile, δ models provide a full suite of diagnostic physical variables and guaranteed mass conservation. Here, we ran experiments to test (1) their ability to extrapolate to regions far from streamflow gauges and (2) their ability to make credible predictions of long-term (decadal-scale) change trends. We evaluated the models based on daily hydrograph metrics (Nash–Sutcliffe model efficiency coefficient, etc.) and predicted decadal streamflow trends. For prediction in ungauged basins (PUB; randomly sampled ungauged basins representing spatial interpolation), δ models either approached or surpassed the performance of LSTM in daily hydrograph metrics, depending on the meteorological forcing data used. They presented a comparable trend performance to LSTM for annual mean flow and high flow but worse trends for low flow. For prediction in ungauged regions (PUR; regional holdout test representing spatial extrapolation in a highly data-sparse scenario), δ models surpassed LSTM in daily hydrograph metrics, and their advantages in mean and high flow trends became prominent. In addition, an untrained variable, evapotranspiration, retained good seasonality even for extrapolated cases. The δ models' deep-network-based parameterization pipeline produced parameter fields that maintain remarkably stable spatial patterns even in highly data-scarce scenarios, which explains their robustness. Combined with their interpretability and ability to assimilate multi-source observations, the δ models are strong candidates for regional and global-scale hydrologic simulations and climate change impact assessment.

54 ENVIRONMENTAL SCIENCES↗

Extended Lagrangian Born–Oppenheimer molecular dynamics for orbital-free density-functional theory and polarizable charge equilibration models

We report extended Lagrangian Born–Oppenheimer molecular dynamics (XL-BOMD) is formulated for orbital-free Hohenberg–Kohn density-functional theory and for charge equilibration and polarizable force-field models that can be derived from the same orbital-free framework. The purpose is to introduce the most recent features of orbital-based XL-BOMD to molecular dynamics simulations based on charge equilibration and polarizable force-field models. These features include a metric tensor generalization of the extended harmonic potential, preconditioners, and the ability to use only a single Coulomb summation to determine the fully equilibrated charges and the interatomic forces in each time step for the shadow Born–Oppenheimer potential energy surface. The orbital-free formulation has a charge-dependent, short-range energy term that is separate from long-range Coulomb interactions. This enables local parameterizations of the short-range energy term, while the long-range electrostatic interactions can be treated separately. The theory is illustrated for molecular dynamics simulations of an atomistic system described by a charge equilibration model with periodic boundary conditions. The system of linear equations that determines the equilibrated charges and the forces is diagonal, and only a single Ewald summation is needed in each time step. The simulations exhibit the same features in accuracy, convergence, and stability as are expected from orbital-based XL-BOMD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗