Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Newton versus the machine: solving the chaotic three-body problem using deep neural networks

ABSTRACT Since its formulation by Sir Isaac Newton, the problem of solving the equations of motion for three bodies under their own gravitational force has remained practically unsolved. Currently, the solution for a given initialization can only be found by performing laborious iterative calculations that have unpredictable and potentially infinite computational cost, due to the system’s chaotic nature. We show that an ensemble of converged solutions for the planar chaotic three-body problem obtained using an arbitrarily precise numerical integrator can be used to train a deep artificial neural network (ANN) that, over a bounded time interval, provides accurate solutions at a fixed computational cost and up to 100 million times faster than the numerical integrator. In addition, we demonstrate the importance of training an ANN using converged solutions from an arbitrary precise integrator, relative to solutions computed by a conventional fixed precision integrator, which can introduce errors in the training data, due to numerical round-off and time discretization, that are learned by the ANN. Our results provide evidence that, for computationally challenging regions of phase space, a trained ANN can replace existing numerical solvers, enabling fast and scalable simulations of many-body systems to shed light on outstanding phenomena such as the formation of black hole binary systems or the origin of the core collapse in dense star clusters.

Breen, Philip G.↗

The Power of Many: An Ensemble Approach to Spectral Similarity

Quantifying the similarity between two mass spectra─a known reference mass spectrum and an unidentified sample mass spectrum─is at the heart of compound identification workflows in gas chromatography–mass spectrometry (GC-MS). The reference spectrum most like the sample is assigned as its identification (provided some quantitative similarity threshold is met, e.g., 80%) and thus accurately measuring similarity is essential. Significant research has gone toward developing metrics for this purpose, each of which has attempted to improve upon existing methods by incorporating GC-MS-specific information (e.g., peak ratios or retention times) or adopting various statistical and algorithmic frameworks. While this active development has led to a plethora of similarity metrics with demonstrated value across different contexts, the unfortunate consequence has been confusion surrounding which metric should be used as a global standard. No such metric is currently accepted as the standard method because different metrics have demonstrated optimal performance in different contexts. In this work, we propose an ensemble approach to spectral similarity scoring that combines the collective information from across existing similarity metrics to form an improved, globally representative similarity metric as a step toward establishing a global standard method. In conclusion, the resulting ensemble metrics are evaluated on over 88,000 spectra of varying complexity and demonstrate improved abilities to accurately rank the correct reference spectrum as the top-matching candidate for a sample relative to the rankings generated by individual similarity scores.

Carbohydrates↗

RCEMIP‐ACI: Aerosol‐Cloud Interactions in a Multimodel Ensemble of Radiative‐Convective Equilibrium Simulations

Aerosol‐cloud interactions are a persistent source of uncertainty in climate research. This study presents findings from a model intercomparison project examining the impact of aerosols on clouds and climate in convection‐permitting radiative‐convective equilibrium (RCE) simulations. Specifically, 11 different modeling teams conducted RCE simulations under varying aerosol concentrations, domain configurations, and sea surface temperatures (SSTs). We analyze the response of domain‐mean cloud and radiative properties to imposed aerosol concentrations across different SSTs. Additionally, we explore the potential impact of aerosols on convective aggregation and large‐scale circulation in large‐domain simulations. The results reveal that the cloud and radiative responses to aerosols vary substantially across models. However, a common trend across models, SSTs, and domain configurations is that increased aerosol loading tends to suppress warm rain formation, enhance cloud water content in the mid‐troposphere, and consequently increase mid‐tropospheric humidity and upper‐tropospheric temperature, thereby impacting static stability. The warming of the upper troposphere can be attributed to reduced lateral entrainment effects due to the higher environmental humidity in the mid‐troposphere. However, models do not agree on aerosol impacts on convective updraft velocity based on the preliminary examination of high‐percentiles of vertical velocity at a single mid‐troposheric layer (500 hPa). In large‐domain simulations, where convection tends to self‐organize, aerosol loading does not consistently influence self‐organization but tends to reduce the intensity of large‐scale circulation forming between convective clusters and dry regions. This reduction in circulation intensity can be explained by the increase in static stability due to the upper tropospheric warming.

54 ENVIRONMENTAL SCIENCES↗

The evolution of voids in the adhesion approximation

We apply the adhesion approximation to study the formation and evolution of voids in the universe. Our simulations-carried out using 128(exp 3) particles in a cubical box with side 128 Mpc-indicate that the void spectrum evolves with time and that the mean void size in the standard Cosmic Background Explorer Satellite (COBE)-normalized cold dark matter (CDM) model with H(sub 50) = 1 scals approximately as bar D(z) = bar D(sub zero)/(1+2)(exp 1/2), where bar D(sub zero) approximately = 10.5 Mpc. Interestingly, we find a strong correlation between the sizes of voids and the value of the primordial gravitational potential at void centers. This observation could in principle, pave the way toward reconstructing the form of the primordialpotential from a knowledge of the observed void spectrum. Studying the void spectrum at different cosmological epochs, for spectra with a built in k-space cutoff we find that the number of voids in a representative volume evolves with time. The mean number of voids first increases until a maximum value is reached (indicating that the formation of cellular structure is complete), and then begins to decrease as clumps and filaments erge leading to hierarchical clustering and the subsequent elimination of small voids. The cosmological epoch characterizing the completion of cellular structure occurs when the length scale going nonlinear approaches the mean distance between peaks of the gravitaional potential. A central result of this paper is that voids can be populated by substructure such as mini-sheets and filaments, which run through voids. The number of such mini-pancakes that pass through a given void can be measured by the genus characteristic of an individual void which is an indicator of the topology of a given void in intial (Lagrangian) space. Large voids have on an average a larger measure than smaller voids indicating more substructure within larger voids relative to smaller ones. We find that the topology of individual voids is strongly epoch dependent, with void topologies generally simplifying with time. This means that as voids grow older they become progressively more empty and have less structure within them. We evaluate the genus measure both for individual voids as well as for the entire ensemble of voids predicted by CDM model. As a result we find that the topology of voids when taken together with the void spectrum is a very useful statistical indicator of the evolution of the structure of the universe on large scales.

Sahni, Varun↗

Thermal Weight Determination and Interstate Coupling in State-Averaged ADAPT-VQE

Characterizing electronic thermal states at low temperatures is an important but challenging task in quantum chemistry and condensed matter physics, making it a prime candidate for a useful application in quantum computing. One of the most successful methods for state preparation on quantum computers is the Adaptive, Problem-Tailored (ADAPT) Variational Quantum Eigensolver (VQE), which has recently been generalized to treat excited states within a state-averaged framework as well as Gibbs states. In this work, we introduce Helmholtz-Optimized Thermal (HOT) ADAPT-VQE, an ancilla-free strategy for preparing Gibbs states that directly minimizes the Helmholtz free energy by targeting the dominant eigenstates of the thermal ensemble. We demonstrate the usefulness of HOT-ADAPT-VQE by predicting the free energy of two model systems with strongly correlated ground states: (1) the Fe 2+ cation in a magnetic field and (2) a [Cu 2 O 7 ] 10– fragment of the Mott insulator La 2 CuO 4 . Our results demonstrate that HOT-ADAPT-VQE significantly improves upon Gibbs-state estimates from multistate variants of ADAPT-VQE, often with substantially shallower quantum circuits, making it a promising candidate for thermal-state calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mars Ground Level Enhancements in the Context of the Solar Energetic Particle Clock

In this work we discuss the growing ensemble of solar particle events registered on the Martian surface, including their temporal appearance and solar sources. Solar energetic particle events from the surface of Mars have been observed starting soon after the August 2012 landing of the Radiation Assessment Detector onboard Curiosity. The Martian atmosphere prevents protons and heavy ions of up to 180 MeV/n kinetic energy from directly reaching the Martian surface. This cut-off is high enough to limit the number of solar energetic particle events measured on the surface to only 15 in ~12 ½ years. Yet we find in this analysis that Mars ground level enhancements follow the distinct SEP clock pattern as proton events observed at lower energies as reported in Posner, Richardson and Strauss (2024). Proton acceleration occurs predominantly near the solar surface, while transport to Mars incurs a delay in onset, and, as we show here, peak, that is a function of the longitudinal magnetic connection distance. between the foot point of solar wind magnetic field lines that intersect the Mars environment with the source longitude of the solar magnetic eruption. A distinct clustering of relative solar source locations at or near the Mars foot points at the Sun’s western limb is apparent, indicating lower flux thresholds from such preferred locations. Our findings have implications for astronaut safety at Mars.

Mars Ground Level Enhancements↗

A Taxonomy-Based Approach to Shed Light on the Babel of Mathematical Models for Rice Simulation

For most biophysical domains, differences in model structures are seldom quantified. Here, we used a taxonomy-based approach to characterise thirteen rice models. Classification keys and binary attributes for each key were identified, and models were categorised into five clusters using a binary similarity measure and the unweighted pair-group method with arithmetic mean. Principal component analysis was performed on model outputs at four sites. Results indicated that (i) differences in structure often resulted in similar predictions and (ii) similar structures can lead to large differences in model outputs. User subjectivity during calibration may have hidden expected relationships between model structure and behaviour. This explanation, if confirmed, highlights the need for shared protocols to reduce the degrees of freedom during calibration, and to limit, in turn, the risk that user subjectivity influences model performance.

model parameterisation↗

Clouds and Convective Self-Aggregation in a Multi-Model Ensemble of Radiative-Convective Equilibrium Simulations

The Radiative-Convective Equilibrium Model Intercomparison Project (RCEMIP) is an intercomparison of multiple types of numerical models configured in radiative-convective56equilibrium (RCE). RCE is an idealization of the tropical atmosphere that has long been used to study basic questions in climate science. Here, we employ RCE to investigate the role that clouds and convective activity play in determining cloud feedbacks, climatecsensitivity, the state of convective aggregation, and the equilibrium climate. RCEMIP is unique amongst intercomparisons in its inclusion of a wide range of model types, including atmospheric general circulation models (GCMs), single column models (SCMs), cloud-resolving models (CRMs), large eddy simulations (LES), and global cloud-resolving models (GCRMs). The first results are presented from the RCEMIP ensemble of more than 30 models. While there are large differences across the RCEMIP ensemble in the representation of mean profiles of temperature, humidity, and cloudiness, in a majority of models anvil clouds rise, warm, and decrease in area coverage in response to an increase in sea surface temperature (SST). Nearly all models exhibit self-aggregation in large domains and agree that self-aggregation acts to dry and warm the troposphere, reduce high cloudiness, and increase cooling to space. The degree of self-aggregation exhibits no clear tendency with warming. There is a wide range of climate sensitivities, but models with parameterized convection tend to have lower climate sensitivities than models with explicit convection. In models with parameterized convection, aggregated simulations have lower climate sensitivities than un-aggregated simulations. Plain Language Summary This study investigates tropical clouds and climate using results from more than 30 different numerical models set up in a simplified framework. The dataset of model simulations is unique in that it includes a wide range of model types configured in a consistent manner. We address some of the biggest open questions in climate science, including how cloud properties change with warming and the role that the tendency of clouds to form clusters plays in determining the average climate and how climate changes. While there are large differences in how the different models simulate average temperature, humidity, and cloudiness, in a majority of models, the amount of high clouds decreases as climate warms. Nearly all models simulate a tendency for clouds to cluster together. There is agreement that when the clouds are clustered, the atmosphere is drier with fewer clouds overall. We don’t find a conclusive result for how cloud clustering changes as the climate warms.

Allison A. Wing↗

The Surface Chemistry of Methanol on Cu 3 Pd(111): Effects of Metal Alloying and Reaction with Hydrogen

Synchrotron-based ambient-pressure X-ray photoelectron spectroscopy (AP-XPS) was used to study the adsorption and surface chemistry of methanol on a Cu 3 Pd(111) model surface. The composition and morphological properties of the pristine Cu 3 Pd(111) substrate were analyzed using a combination of low-energy electron diffraction (LEED), scanning tunneling microscopy (STM), and low-energy ion scattering (LEIS). The results of ion scattering showed segregation of Pd toward the surface, with a Pd/Cu ratio close to 0.5, a larger value than the ratio of 0.33 expected for a Cu 3 Pd bimetallic structure. The surface of the alloy exhibited a good LEED pattern with three-fold long-range periodicities. In STM, clusters of palladium with hexagonal arrays and Pd- Pd distances of 2.7-2.8 Å were detected. Bonding to copper perturbed the electronic properties of the atoms in the Pd clusters, shifting their 4d states toward higher binding energy with respect to the Fermi level. Further, the valence band spectrum of Cu 3 Pd(111) exhibited a line shape that was very different from those displayed by Cu(111) or Pd(111). At low pressures, the adsorption of methanol on Cu 3 Pd(111) at 300 K mainly produced CH 3 O, CO and CH x species. AP-XPS showed that most Pd atoms in the surface of the Cu 3 Pd(111) alloy interacted with the decomposition products of methanol. No significant changes were observed in the core levels of copper upon the adsorption and dissociation of methanol, suggesting that the molecule mainly interacted with Pd sites of the alloy. Reaction with hydrogen led to fast removal of CH x , C and PdC x species from Cu 3 Pd(111) at moderate (< 450 K) temperatures and prevented a CH x → C transformation. If the stability of adsorbed CH 3 O is used as a descriptor for the hydrogenation of CO 2 to methanol, Cu 3 Pd should be a much better catalyst than monometallic palladium. This may be a consequence of electronic and ensemble effects in the alloy that moderate the reactivity of Pd sites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wave function methods for canonical ensemble thermal averages in correlated many-fermion systems

We present a wave function representation for the canonical ensemble thermal density matrix by projecting the thermofield double state against the desired number of particles. Furthermore, the resulting canonical thermal state obeys an imaginary-time evolution equation. Starting with the mean-field approximation, where the canonical thermal state becomes an antisymmetrized geminal power (AGP) wave function, we explore two different schemes to add correlation: by number-projecting a correlated grand-canonical thermal state and by adding correlation to the number-projected mean-field state. As benchmark examples, we use number-projected configuration interaction and an AGP-based perturbation theory to study the hydrogen molecule in a minimal basis and the six-site Hubbard model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A multi-dimensional parametric study of variability in multi-phase flow dynamics during geologic CO 2 sequestration accelerated with machine learning

Successful geologic CO 2 storage projects depend on numerical simulations to predict reservoir performance during site selection, injection verification, and post-injection monitoring phases of the project. These numerical simulations solve non-linear sets of coupled partial differential equations, while accounting for multi-phase fluid dynamics on the basis of constitutive equations that are embedded into the solution scheme. As a consequence, individual simulations often require tens to hundreds of hours to complete on high-performance computing clusters. Moreover, laboratory experiments reveal that parametric functions for capillary pressure and relative permeability exhibit substantial variability, even within the same rock type. This combination of computational expense and wide-ranging parametric variability means that there remains substantial uncertainty in the behavior of multi-phase CO 2 -water systems, particularly in the context of feedbacks between relative permeability and capillary pressure. To bridge this knowledge gap, here we develop a novel workflow that utilizes physics-based numerical simulation to train an artificial neural network (ANN) emulator for interrogating the multivariate parameter space that governs both capillary pressure and relative permeability. With this approach, the ANN is trained to emulate both fluid pressure distribution and CO 2 saturation, which are then interrogated quantitatively to generate parametric response surface mappings with high-fidelity resolution. Results from this study initially show that capillary entry pressure is the dominant control on both CO 2 plume geometry and fluid pressure propagation when considering the combined effects of capillary pressure and relative permeability, particularly when phase interference is low and residual CO 2 saturation is high. Moreover, the ANN emulator provides tremendous computational speed-up by computing 2691 individual simulations in several minutes; whereas, the same simulation ensemble would have required ~3 years of simulation time using only physics-based simulation methods (25,000 times speed up).

58 GEOSCIENCES↗

Emerging Electrochemical Techniques for Probing Site Behavior in Single-Atom Electrocatalysts

Single-atom catalysts (SACs) have aroused tremendous interest over the past decade, particularly in the community of energy and environment-related electrocatalysis. A rapidly growing number of recent publications have recognized it as a promising candidate with maximum atomic utilization, distinct activity, and selectivity in comparison to bulk catalysts and nanocatalysts. However, the complexity of localized coordination environments and the dispersion of isolated sites lead to significant difficulties when it comes to gaining insight into the intrinsic behavior of electrocatalytic reactions. Furthermore, the low metal loadings of most SACs make conventional ensemble measurements less likely to be accurate on the subnanoscale. Thus, it remains challenging to probe the activity and properties of individual atomic sites by available commercial instruments and analytical methods. In spite of this, continuing efforts have lately focused on the development of advanced measurement methodologies, which are very useful to the fundamental understanding of SACs. There have recently been a number of in situ/operando techniques applied to SACs, such as electron microscopy, spectroscopy, and other analysis methods, which support relevant functions to identify the active sites and reaction intermediates and to investigate the dynamic behavior of localized structures of the catalytic sites. This Account aims to present recent electrochemical probing techniques which can be used to identify single-atomic catalytic sites within solid supports. First, we describe the basic principles of molecular probe methods for the study and analysis of electrocatalytic site behavior. In particular, the in situ probing technique enabled by surface interrogation scanning electrochemical microscopy (SISECM) can measure the active site density and kinetic rate with high resolution. An alternative electrochemical probing technique is further demonstrated on the basis of single-entity electrochemistry, which allows the unique electrochemical imaging of the size and catalytic rate of single atoms, molecules, and clusters. Next, the merits and limitations of different electrochemical techniques are then discussed, along with perspectives for future prospects. Apart from this, we further showcase the powerful capability of emerging electrochemical probing techniques for determining significant effects and properties of SACs for various electrocatalytic reactions, including oxygen reduction and evolution, hydrogen evolution, and nitrate reduction. Overall, electrochemical techniques with atomic resolution have greatly increased opportunities for observing, measuring, and understanding the surface and interface chemistry during energy conversion. In the future, it is anticipated that the development of electrochemical probing techniques will be advanced with innovative perspectives on the behavior and features of SACs. We hope that this Account can contribute in several ways to promoting the fundamental knowledge and technical progress of emerging electrochemical measurements for studying SACs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fast calculation of diffraction patterns from an ensemble of aligned molecules

We report an algorithm to calculate electron diffraction patterns for molecules with anisotropic angular distribution, which is significantly faster than existing methods. The algorithm uses a transform to convert the molecular orientation distribution, which is a function of three Euler angles, to the atom-pair distribution functions which depend on the polar and azimuthal angles. The diffraction signal can then be calculated from the atom-pair distributions. We demonstrate the computation method numerically by calculating electron diffraction patterns for a symmetric top molecule (trifluoroiodomethane) and an asymmetric top molecule (formaldehyde) and show that it reduces the calculation time by approximately two orders of magnitude compared to the standard brute-force method. Here, the method can also be applied to the calculation of x-ray diffraction patterns.

74 ATOMIC AND MOLECULAR PHYSICS↗

Connecting the Radiative Influences of Aerosol upon the Mass Flux Profiles of Shallow Cumuli across the Southeast Atlantic Ocean Basin and its Boundaries (Final Report)

The Atlantic Ocean covers approximately 25% of Earth’s surface and the atmosphere above it is home to a complex array of clouds and aerosols that have important influences on regional and global weather and climate. These influences must be accurately depicted in short range, medium range, seasonal, and climate forecast models. Conditions over the tropical Atlantic are particularly complex due to continental scale plumes of dust from the Sahara Desert and smoke from agricultural burning in Africa that drift across the Atlantic Ocean basin toward the Americas. These plumes are often found meandering in the lower atmosphere above shallow tropical clouds that form above the ocean surface, presumably mingling with these clouds on occasion due to convective mixing processes. Elevated dust and smoke particles absorb incoming sunlight and substantially warm the marine atmosphere in the layer in which they are present. This warming may alter the thermal stability of the marine atmosphere and may throttle or enhance the development of clouds, change their internal structure and the rate at which they precipitate. Alternatively, it may isolate the lower atmosphere from drier layers above enabling water vapor to accumulate near the ocean surface potentially leading to the development of deeper convection. Our study was organized around the principal concept of determining the impact of African smoke and dust plumes upon cloud development at ASI and understanding how a popular shallow convection parameterization used in models responds to the presence of this aerosol. We analyzed observations collected during the US Department of Energy (DOE) Layered Atlantic Smoke Interactions with Clouds (LASIC) campaign using the Atmospheric Radiation Measurement Program’s Mobile Facility #1 (AMF-1), which was deployed to Ascension Island (ASI) in Southeastern Atlantic for a one-year period. Ascension Island is immediately downwind from an African source of these plumes, but far enough removed to enable the lower atmosphere to have reacted to their presence. A main goal of our study was to compute radiative heating rate profiles over Ascension Island and along the trajectory from the biomass burning regions along coastal Africa to Ascension Island. To compute the radiative heating rate due to aerosols and clouds, we employed observed profiles of temperature, humidity, and clouds from LASIC alongside aerosol optical properties from the Modern-era Retrospective analysis for Research and Applications Version 2 (MERRA-2), as input for the Rapid Radiation Transfer Model (RRTM). Radiative heating was also assessed across the southeast Atlantic Ocean using an ensemble of back trajectories from the Hybrid Single Particle Lagrangian Integrated Trajectory (HYSPLIT) model. We were successful in this effort and the resulting publication is already being cited despite its relatively short lifetime in the literature. The second part of our study involved the process-level representation of the clouds observed over the Southeastern Atlantic in models. Our initial task, upon which the balance of this portion of the study depended, was to evaluate the representativeness of the clouds observed by AMF-1. What orographic influence did Ascension Island have on the measured cloud structure? We set out to answer this question by trying to separate observations that were clearly indicative of orographic forcing from those that were representative of open ocean. The siting of AMF-1, which was 340-m above sea-level on the slope of a steep escarpment, proved demonstrably problematic for cloud process measurements. Despite considerable effort using multiple approaches including artificial intelligence (AI), we were unable to successfully compensate for the orographic forcing present at the AMF-1 site on ASI and, hence, produce process-level summaries of the convective mass flux representative of open ocean over the Southeastern Atlantic. Even though the project has officially ended, we are making a final attempt using two new types of clustering algorithm (i.e., AI) on the recommendation of a former student, who is an expert in this area. Since we have only recently begun to test these new AI schemes, the recommendations from the second part of our project outlined below are based on our experience at the time of this report.

54 ENVIRONMENTAL SCIENCES↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Simulation of Radiation-Induced DNA Damage With the Code RITRACKS

INTRODUCTION DNA damage is one of the most physiologically important effects of ionizing radiation. Clustered DNA damage events, like double-strand breaks (DSBs), have the most notable biological consequences. DNA damage types depend on both the track structure of the radiation and the spatial organization of the DNA. High linear energy transfer (LET) charged nuclei, found in galactic cosmic rays (GCR), are known to produce large numbers of complex DNA damage events. The human genome is packaged into chromatin, which can take on locus-dependent and cell type-dependent spatial conformations that correspond to epigenetic states, such as more open, extended structures in transcriptionally active chromatin. These epigenetic differences can affect DNA break patterns in response to ionizing radiation, potentially creating distinct DNA repair and signaling outcomes across the genome in different cells. MATERIAL AND METHODS The code RITRACKS (Relativistic Ion Tracks), which simulates stochastic radiation track structures and radiation chemistry, was used to model damage on isolated and histone-bound DNA by various types of ions and photons. The changes made to the code to perform radiation-induced DNA damage, and simulation results on single nucleosomes are given in our recent paper. In this work, the DNA building capabilities of RITRACKS have been extended to simulate more complex DNA structures build on the coarse-grain simulation framework meso-WLCsim. This code can sample generic chromatin fiber conformation ensembles based on the geometry of nucleosomes and mechanical properties of DNA. Using RITRACKS, we simulated the fragment length distributions (FLD) of irradiated DNA structures built using the chromatin conformations of WLCsim and obtained results representative of those obtained with Radiation-Induced Correlated Cleavage with sequencing (RICC-Seq) experiments [6]. We have also performed Fe ion and photon irradiations of K562, IMR90, BJ and RPE-1 cells at NSRL to experimentally validate results. Sample processing and data analysis are in progress and any available preliminary results will be discussed. DISCUSSION The recent updates in the code RITRACKS allow the calculation of several quantities such as the DNA damage yield and the FLD. This approach can be used to model epigenetic state-specific chromatin structure parameters to leverage the epigenetic state data available for many human cell types to infer relative DNA damage sensitivity among genomic loci.

I Plante↗

Large-scale atomistic model construction of subbituminous and bituminous coals for solvent extraction simulations with reactive molecular dynamics

Large-scale atomistic models for complex polycyclic aromatic hydrocarbon systems help understand the chemical properties and behaviors of complex feedstocks such as coal or petroleum. However, the development and utilization of large-scale models remain limited due to the difficulty in achieving the varied structural characteristics necessary to capture stochastic nature of these feedstocks. Here we demonstrate a systematic workflow to construct stochastic molecular systems from a broad analytical suite: high-resolution transmission electron microscopy (HRTEM), carbon-13 nuclear magnetic resonance spectroscopy ( 13 C NMR), laser desorption ionization mass spectroscopy (LDI-MS), and elemental analysis. We present a model construction and analysis utility of a new Python-based module. We selected one subbituminous and three high-volatile bituminous coals to construct large-scale models (~40,000 atoms). The constructed models were utilized to examine the affinity for solvent extraction (naphthalene or tetralin) and the effect of structural properties (e.g., aromatic cluster size, functional groups, and cross-linking) in reactive molecular dynamics simulations. Complex chemical reactions were monitored with bond order transitions, intermediates formation, and mass distributions. Reactive molecular dynamics simulations suggest a plausible chemical extraction process and products for the complex fossil feedstocks. The results indicated that radical formations with bond breaking of bridging oxygens and carbons were required at high temperatures to facilitate hydrogeneration and extraction of gas molecules from radical-free molecules. We observed that aliphatic chains of tetralin were easily decomposed and combined with radicals to form small size of molecules with aryl bonding, mainly increasing molecules in the 500–1000 Da, while naphthalene had little impact on chemical extraction process.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael↗