Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large-scale data association”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Federated learning for 2D synchrotron x-ray diffractometry: a cross-institutional approach for phase quantification of Ti–6Al–4V alloy

High-energy Two dimensional (2D) synchrotron x-ray diffractometry provides important insights into the atomistic structure and phase evolution of materials, yet traditional analysis methods remain complex, knowledge-intensive, and computationally demanding. Deep-learning models offer a powerful alternative for automating their analysis. Institutions that hold these datasets may be unwilling to share their data due to privacy and security policies, as well as the challenges associated with large-scale data transfer. As a result, models trained on local datasets often perform well only on their own data but exhibit bias and poor generalization across different instruments or facilities. To overcome these limitations, we explore federated learning (FL) for 2D synchrotron diffractograms, enabling collaborative model training without exchanging raw data. In this study, 2D synchrotron diffractograms of Ti–6Al–4V alloy collected from two independent facilities are used to train convolutional neural networks for predicting the β-phase volume fraction. Experimental results show that federated global models significantly outperform locally trained models in terms of generalization and achieve accuracy comparable to centralized trained models. These findings demonstrate the potential of FL to enable secure, cross-institutional collaboration and enhance the scalability of deep-learning-based materials characterization.

36 MATERIALS SCIENCE↗

Bridging the Gap for Powering Data Centers

The rapid expansion of data centers, primarily driven by artificial intelligence, is outpacing the adaptability of the U.S. electric grid. This report, developed by Idaho National Laboratory (INL) , presents a gap analysis of some of the infrastructure challenges associated with large-scale data center deployment. Drawing from a national workshop hosted by INL in October of 2025, the report synthesizes stakeholder insights, survey data, and technical discussions to identify critical barriers and research needs. Key findings highlight the growing preference for behind-the-meter generation, the perceived inadequacy of legacy interconnection processes, and the urgent need for improved coordination between utilities, regulators, and data center developers. Environmental concerns such as water use and noise pollution, as well as economic constraints like equipment lead times and cost allocation, are also explored. The report outlines national lab capabilities in modeling, simulation, and technical assistance, and proposes targeted R&D priorities to support resilient, scalable, and efficient integration of data centers into the grid.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Seismic Elastic Double-Beam Characterization of Faults and Fractures for CO₂ Storage Site Selection

Site characterization for underground injection and storage of gigatonne-scale CO₂ requires reliable and cost-effective methods to detect and characterize faults and fractures and to assess their stress state and fault activation potential. This is critical, as wastewater injection and disposal have been shown to activate faults and induce earthquakes, and CO₂ leakage remains a key concern for long-term storage. In this project, we developed seismic methods to detect and characterize large-scale sedimentary and crystalline basement faults and associated small-scale fractures below conventional seismic imaging resolution using multicomponent (9C) surface seismic data. Machine learning was used to automatically interpret large-scale faults, providing key information for estimating the maximum magnitude of potential induced earthquakes. High-fidelity imaging was achieved by exploiting redundancy across multiple elastic wave modes, where independent images from different modes and frequencies cross-validate each other. We also used our nonlinear signal comparison (NLSC) method for ground roll removal, improving data quality in complex near-surface conditions. The methods were validated using field data acquired in central Montana. Results show that basement faults extend into the sedimentary section and that small-scale fractures are widespread above the basement. The inferred stress orientation is consistent with regional stress data, and the estimated maximum induced earthquake magnitude is small (Mw ~2.3). The developed workflow provides a practical approach for fault and fracture characterization and for assessing induced seismicity and leakage risk. It is directly applicable to CO₂ storage site selection and to other subsurface systems.

02 PETROLEUM↗

Advanced Mooring Lines for Marine Energy Workshop – Summary Report

This report summarizes the outcomes of the virtual Advanced Mooring Lines for Marine Energy Workshop. The workshop convened experts from industry, academia, and national laboratories to identify key challenges associated with the deployment of synthetic mooring lines for marine energy applications. Discussions focused on four technical areas—materials, failure and degradation, testing, and modeling. Despite the significant promise of leveraging synthetic mooring systems for marine energy applications, discussion highlighted broader research and development needed to address the lack of validated long-term performance data, insufficient structural health monitoring frameworks, limited large-scale testing capabilities, and gaps in predictive modeling. Overall, the workshop emphasized the importance of coordinated research and development efforts to advance reliable, durable, and cost-effective synthetic mooring solutions for marine energy applications.

16 TIDAL AND WAVE POWER↗

Brochure on the 2024 ASCR Workshop on Energy-Efficient Computing for Science

Large-scale computing has enabled numerous scientific discoveries, including ground-breaking achievements facilitated by the US Department of Energy (DOE) supercomputers and advances in applied mathematics and computer science. While important advances were made in energy efficiency to enable exascale computing, continued efforts are needed to dramatically improve the energy efficiency of the next generation of high-performance computing (HPC) systems and, more broadly, AI data centers. Without substantial improvements in energy efficiency, the energy consumption associated with computing could become a limiting factor for future scientific discovery, national security, and technological advancement.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Molecular origin of negative lithium transference in electrolytes with star-shaped multivalent anions

Large-scale molecular dynamics simulations illustrate that highly correlated cation–anion motion leads to negative t 0+ on the order of 1 in lithium electrolytes with star-shaped multivalent anions. Large multivalent anions have gained increasing attention for their potential to improve lithium transference in electrolytes. We employ large-scale molecular dynamics simulations based on the Onsager transport framework to investigate ion transport in a lithium electrolyte with star-shaped multivalent anions. The simulations show that t 0+, the cation transference number with respect to solvent velocity, is negative over a wide range of concentration. This is consistent with experimental data reported previously. The simulation-based Onsager transport coefficients reveal that the magnitudes of the cation–cation, anion–anion, and cation–anion correlations are comparable, a signature of highly correlated motion in the electrolyte. Examination of the cation solvation environment indicates the presence of strong cation–anion association across the entire concentration range, which leads to negative t 0+ on the order of −1. Both simulation and experiment also show that the maximum value of t 0+ reaches 0 when the cation concentration is c + = 0.4 M. This is the concentration at which the anions begin to spatially overlap, and lithium ions serve as dynamic linkers to balance cation–cation and cation–anion correlations. Our results provide molecular-level insights into the origin of transference in multivalent electrolytes.

Fang, Chao↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

Taylor approximation variance reduction for approximation errors in PDE-constrained Bayesian inverse problems

In numerous applications, surrogate models are used as a replacement for accurate parameter-to-observable mappings when solving large-scale inverse problems governed by partial differential equations (PDEs). The surrogate model may be a computationally cheaper alternative to the accurate parameter-to-observable mappings and/or may ignore additional unknowns or sources of uncertainty. The Bayesian approximation error (BAE) approach provides a means to account for the induced uncertainties and approximation errors, i.e. the errors between the accurate parameter-to-observable mapping and the surrogate. The statistics of these errors are, however, in general unknown a priori, and are thus calculated using Monte Carlo sampling. Although the sampling is typically carried out offline, i.e. before considering the data, the process can still represent a computational bottleneck. In this work, we develop a scalable computational approach for reducing the costs associated with the sampling stage of the BAE approach. Specifically, we consider the Taylor expansion of the accurate and surrogate forward models with respect to the uncertain parameter fields either as a control variate for variance reduction or as a means to directly and efficiently approximate the mean and covariance of the approximation errors. We propose efficient methods for evaluating the expressions for the mean and covariance of the Taylor approximations based on linear(-ized) PDE solves. Furthermore, the proposed approach is independent of the dimension of the uncertain parameter, depending instead on the intrinsic dimension of the data, ensuring scalability to high-dimensional problems. The potential benefits of the proposed approach are demonstrated for two high-dimensional inverse problems governed by PDE examples, namely for the estimation of a distributed Robin boundary coefficient in a linear diffusion problem, and for a coefficient estimation problem governed by a nonlinear diffusion problem.

Bayesian approximation error↗

ReMU: regional minimal updating for model-based derivative-free optimization

Derivative-free optimization (DFO) problems are optimization problems where derivative information is unavailable or extremely difficult to obtain. Model-based DFO solvers have been applied extensively in scientific computing. Powell's NEWUOA (2004) [Powell, The NEWUOA software for unconstrained optimization without derivatives, in Large-Scale Nonlinear Optimization, Nonconvex Optimization and its Applications Vol. 83, G. Di Pillo and M. Roma, eds., Springer, 2006, pp. 255–297] and Wild's POUNDerS (2014) [Wild, Solving derivative-free nonlinear least squares problems with POUNDERS, in Advances and Trends in Optimization with Engineering Applications, T. Terlaky, M.F. Anjos, and S. Ahmed, eds., SIAM, 2017, pp. 529–540] explore the numerical power of the minimal norm Hessian (MNH) model for DFO and contributed to the open discussion on building better models with fewer data to achieve faster numerical convergence. Another decade later, we propose the regional minimal updating (ReMU) models, and extend the previous models into a broader class, including the H 2 norm models [Xie and Yuan, Least H 2 norm updating of quadratic interpolation models for derivative-free trust-region algorithms, IMA J. Numer. Anal. 46 (2025), pp. 21–50]. This paper shows motivation behind ReMU models, computational details, theoretical and numerical results on particular extreme points and the barycentre of ReMU's weight coefficient region, and the associated KKT matrix error and distance. Novel metrics, such as the truncated Newton step error, are proposed to numerically understand the new models' properties. A new algorithmic strategy, based on iteratively adjusting the ReMU model type, is also proposed, and shows numerical advantages by combining and switching between the barycentric model and the classic least Frobenius norm model in an online fashion.

derivative-free trust-region methods↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Discovery of a nearby radio relic in the low-mass, merging cluster Abell 4067

Shock waves generated during cluster mergers offer a powerful probe of how large-scale structure grows and evolves in the Universe. As part of the MeerKAT-South Pole Telescope (SPT) survey, we report the discovery of a single arc-like radio relic in the galaxy cluster Abell 4067 ($z=0.099$), one of the lowest-mass clusters known to host such a structure. MeerKAT UHF-band (0.58–1.09 GHz) observations reveal a relic with a largest linear size of $\sim 1.48\pm 0.02$ Mpc, located at a projected distance of 0.95 Mpc from the cluster centre. XMM–Newton X-ray data show that the relic’s position and orientation relative to the intracluster medium (ICM) elongation are consistent with a merger-driven shock-wave scenario. The relic has an estimated radio power of $3.10\pm 0.03\times 10^{24}$ W Hz$^{-1}$ at 150 MHz. When placed in the $P_{150\, \mathrm{MHz}}$–$M_{500}$ scaling relation, the Abell 4067 relic appears less luminous compared to relics in more massive clusters, suggesting an association with weak merger shocks. This finding supports the idea that relics in low-mass clusters may form through less energetic merger events, leading to weak merger shocks. The latter is supported by the absence of a detectable central radio halo in Abell 4067, which reinforces the idea that luminous radio haloes are not a universal outcome of cluster mergers and highlights the role of cluster mass, merger energetics, and evolutionary stage in shaping diffuse radio emission in the ICM.

79 ASTRONOMY AND ASTROPHYSICS↗

Proposed New STELLA-2 Tests

In 2025, KAERI and Argonne initiated a collaboration based on software validation using data from tests performed at KAERI’s STELLA-2 facility. STELLA-2 is a large-scale sodium thermal-hydraulic test facility that was originally developed to support KAERI’s development of the PGSFR design concept. KAERI has performed a large number of tests at STELLA-2, providing valuable data to support sodium fast reactor code validation. KAERI will be conducting new tests in 2026, targeting system conditions that were not achieved during previous test campaigns. KAERI has requested from Argonne proposed tests that could be performed during the upcoming test campaign. This report documents Argonne’s proposed tests in fulfillment of KAERI’s request.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

ARCTRON: A Rapid Experimental Proving Ground for TPS Experiments and Arcjet Technology Development

Innovation in high-enthalpy facilities is fundamentally limited by the cost and risk of experimentation. New concepts for plasma control, diagnostics, facility components, and plasma-material interaction often require repeated iterations that are impractical to perform in production arcjets. As a result, promising ideas may remain unexplored or reach operational facilities only after significant development effort. ARCTRON is being developed as a rapid experimental proving ground where new ideas in plasma science, arcjet engineering, diagnostics, and material response can be conceived, tested, and quantitatively evaluated before transition to large-scale facilities. The platform combines radio-frequency (RF) and DC arc plasma generation, externally applied magnetic fields, configurable gas composition, reduced-pressure operation, laser heating, electrical biasing, and modular diagnostic access. These capabilities permit the plasma source, applied forcing, test article, and measurement configuration to be modified independently, allowing individual physical mechanisms to be isolated more readily than in a traditional test environment. One class of investigations addresses fundamental plasma-surface interaction physics. Conventional material tests often expose a specimen simultaneously to convective heating, reactive species, pressure, shear, radiation, and surface-current effects. The resulting material response may be measured accurately, while the contribution of each mechanism remains difficult to identify. ARCTRON is designed to vary these effects selectively. Plasma chemistry can be changed independently through configurable gas mixtures; magnetic fields and electrical biasing can modify charged-particle transport; laser heating can provide a non-plasma thermal input; and pressure, flow, and discharge mode can be varied over a broad operating space. This enables controlled tests of hypotheses involving surface catalycity, reactive-species transport, plasma-assisted oxidation, electromagnetic effects, shear, and the relative contributions of thermal and chemical loading. A second class of investigations enabled by this approach concerns the engineering of high-enthalpy facilities themselves. Arc-heated facilities are limited by electrode erosion, unstable arc attachment, localized heating, and damage to nozzles and other plasma-facing components. ARCTRON provides a lower-cost environment for testing concepts intended to mitigate these limitations. Candidate investigations include the use of applied magnetic fields to alter current paths and reduce plasma interaction with nozzle walls, ExB forcing to introduce controlled plasma rotation, magnetic or geometric approaches for distributing arc attachment, and alternative electrode or discharge configurations intended to reduce erosion and improve stability. Because the platform is reconfigurable, these concepts can be evaluated through repeated design--build--test cycles before they are considered for implementation in operational facilities. The platform also supports the development and validation of diagnostics that may be difficult to introduce initially into a large arcjet. Current and planned measurements include spatially resolved optical emission spectroscopy, electrostatic probes, fast imaging, pyrometry, calorimetry, laser-induced fluorescence, and absorption spectroscopy. These diagnostics are intended not merely to document a nominal operating condition, but to constrain the local plasma state and its relationship to component or material response. The modular facility geometry allows diagnostic concepts to be tested, calibrated, and compared under repeatable conditions before deployment in more demanding environments. ARCTRON is also supported by an integrated software suite. Automated control and data acquisition allow discharge parameters, gas composition, magnetic fields, diagnostic timing, and test configuration to be recorded as part of each experiment (STARDAC - Software for Testing, Analysis, Research Data, and Control). The Backend for Experiment Analysis, Storage, and Traceability (BEAST) is a database that provides the infrastructure needed to associate heterogeneous measurements with facility configuration, specimen identity, calibration state, geometry, and analysis provenance. This backend is particularly important for exploratory campaigns, in which many related configurations may be tested, and the value of an individual experiment depends on its connection to earlier and subsequent iterations. Complementary analysis capabilities, including computer-vision-based transient response measurements (arcjetCV), three-dimensional surface reconstruction (STARSCAN), and model-based Bayesian inference (SHIELD), and tomography data analysis (TOMATO, PuMA) can be incorporated when required by a specific hypothesis without becoming the focus of every campaign. The central objective of ARCTRON is therefore not to maximize heat flux or reproduce a complete flight environment. Its purpose is to reduce the cost and time required to ask consequential questions about plasma behavior, plasma-facing materials, diagnostics, and arcjet technology. By providing a controlled environment for rapid reconfiguration, mechanism isolation, quantitative measurement, and iterative engineering, ARCTRON can help mature concepts that would otherwise remain too speculative or too risky for evaluation in production facilities. The resulting knowledge can then guide the design of material models, focus test objectives in larger arcjets, reduce facility-development risk, and improve the physical basis of high-enthalpy ground testing. This work will present the ARCTRON architecture, operating modes, diagnostic suite, and digital experimental workflow. Initial experimental results from the first integrated operation of the facility will be presented, including flow characterization, power limitations, and deployment of the initial diagnostic suite. Ongoing development efforts aimed at catalycity characterization, magnetic plasma control, and advanced optical diagnostics will also be discussed, illustrating how the platform supports rapid iteration from concept to experiment.

experimental diagnostics↗

ARCTRON: A Rapid Experimental Proving Ground for TPS Experiments and Arcjet Technology Development

Innovation in high-enthalpy facilities is fundamentally limited by the cost and risk of experimentation. New concepts for plasma control, diagnostics, facility components, and plasma-material interaction often require repeated iterations that are impractical to perform in production arcjets. As a result, promising ideas may remain unexplored or reach operational facilities only after significant development effort. ARCTRON is being developed as a rapid experimental proving ground where new ideas in plasma science, arcjet engineering, diagnostics, and material response can be conceived, tested, and quantitatively evaluated before transition to large-scale facilities. The platform combines radio-frequency (RF) and DC arc plasma generation, externally applied magnetic fields, configurable gas composition, reduced-pressure operation, laser heating, electrical biasing, and modular diagnostic access. These capabilities permit the plasma source, applied forcing, test article, and measurement configuration to be modified independently, allowing individual physical mechanisms to be isolated more readily than in a traditional test environment. One class of investigations addresses fundamental plasma-surface interaction physics. Conventional material tests often expose a specimen simultaneously to convective heating, reactive species, pressure, shear, radiation, and surface-current effects. The resulting material response may be measured accurately, while the contribution of each mechanism remains difficult to identify. ARCTRON is designed to vary these effects selectively. Plasma chemistry can be changed independently through configurable gas mixtures; magnetic fields and electrical biasing can modify charged-particle transport; laser heating can provide a non-plasma thermal input; and pressure, flow, and discharge mode can be varied over a broad operating space. This enables controlled tests of hypotheses involving surface catalycity, reactive-species transport, plasma-assisted oxidation, electromagnetic effects, shear, and the relative contributions of thermal and chemical loading. A second class of investigations enabled by this approach concerns the engineering of high-enthalpy facilities themselves. Arc-heated facilities are limited by electrode erosion, unstable arc attachment, localized heating, and damage to nozzles and other plasma-facing components. ARCTRON provides a lower-cost environment for testing concepts intended to mitigate these limitations. Candidate investigations include the use of applied magnetic fields to alter current paths and reduce plasma interaction with nozzle walls, ExB forcing to introduce controlled plasma rotation, magnetic or geometric approaches for distributing arc attachment, and alternative electrode or discharge configurations intended to reduce erosion and improve stability. Because the platform is reconfigurable, these concepts can be evaluated through repeated design--build--test cycles before they are considered for implementation in operational facilities. The platform also supports the development and validation of diagnostics that may be difficult to introduce initially into a large arcjet. Current and planned measurements include spatially resolved optical emission spectroscopy, electrostatic probes, fast imaging, pyrometry, calorimetry, laser-induced fluorescence, and absorption spectroscopy. These diagnostics are intended not merely to document a nominal operating condition, but to constrain the local plasma state and its relationship to component or material response. The modular facility geometry allows diagnostic concepts to be tested, calibrated, and compared under repeatable conditions before deployment in more demanding environments. ARCTRON is also supported by an integrated software suite. Automated control and data acquisition allow discharge parameters, gas composition, magnetic fields, diagnostic timing, and test configuration to be recorded as part of each experiment (STARDAC - Software for Testing, Analysis, Research Data, and Control). The Backend for Experiment Analysis, Storage, and Traceability (BEAST) is a database that provides the infrastructure needed to associate heterogeneous measurements with facility configuration, specimen identity, calibration state, geometry, and analysis provenance. This backend is particularly important for exploratory campaigns, in which many related configurations may be tested, and the value of an individual experiment depends on its connection to earlier and subsequent iterations. Complementary analysis capabilities, including computer-vision-based transient response measurements (arcjetCV), three-dimensional surface reconstruction (STARSCAN), and model-based Bayesian inference (SHIELD), and tomography data analysis (TOMATO, PuMA) can be incorporated when required by a specific hypothesis without becoming the focus of every campaign. The central objective of ARCTRON is therefore not to maximize heat flux or reproduce a complete flight environment. Its purpose is to reduce the cost and time required to ask consequential questions about plasma behavior, plasma-facing materials, diagnostics, and arcjet technology. By providing a controlled environment for rapid reconfiguration, mechanism isolation, quantitative measurement, and iterative engineering, ARCTRON can help mature concepts that would otherwise remain too speculative or too risky for evaluation in production facilities. The resulting knowledge can then guide the design of material models, focus test objectives in larger arcjets, reduce facility-development risk, and improve the physical basis of high-enthalpy ground testing. This work will present the ARCTRON architecture, operating modes, diagnostic suite, and digital experimental workflow. Initial experimental results from the first integrated operation of the facility will be presented, including flow characterization, power limitations, and deployment of the initial diagnostic suite. Ongoing development efforts aimed at catalycity characterization, magnetic plasma control, and advanced optical diagnostics will also be discussed, illustrating how the platform supports rapid iteration from concept to experiment.

experimental diagnostics↗