Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed average tracking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

45 records · Page 3

Test Beam Results of Planar Pixel Sensor for the CMS Phase 2 Inner Tracker Upgrade

Results of the test beam measurements that characterise the performance of CMS Readout Chip (CROC) sensors to be used in the High Luminosity era of the Large Hadron Collider (HL-LHC) are presented. The HL-LHC peak instantaneous luminosity of $7.5 \times 10^{34} \ \text{cm}^{-2} s^{-1}$ corresponds to an average of around 200 inelastic proton-proton collisions per beam-crossing every 25 ns. In order to efficiently reconstruct and track particles in this extreme and challenging conditions, the present CMS tracking detector will be completely replaced. The new tracking detector consists of an Inner Tracker closest to the beamline and an Outer Tracker surrounding it. These are populated with modules that comprise of readout chips and silicon sensors. The test beam measurements of these modules are vital to understand the performance of the related technologies. Using a primary 120 GeV proton beam from the Main Injector at Fermilab, data was collected at the Fermilab Test Beam Facility (FTBF) using the silicon tracker telescope that provides a precision position measurement of the track impact point with less than 5 $\mu$m uncertainty. The proton beam was incident on a 1x2 planar CROC module developed by Hamamatsu. The sensor has 100x25 $\mu m^2$ standard pixels and also a smaller number of 225x25 $\mu m^2$ longer pixels at the boundary between the two ROCs. We present characterization of these modules that includes pixel efficiency, resolution, cluster size and charge distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

In situ measurement of three-dimensional intergranular stress localizations and grain yielding under elastoplastic axial-torsional loading

The three-dimensional grain-averaged response of solid bar samples under non-proportional (NP) elastoplastic axial-torsional loading was investigated using in situ high energy diffraction microscopy (HEDM) and companion crystal plasticity finite element (CPFE) modeling. Important stress metrics including applied shear (σ θZ ) and axial (σ ZZ ) stress tensor components, stress and stress deviator tensor invariants (I 1 , J 2 , and J 3 ), von Mises equivalent stress (σ$^{grain}_{VM}$), maximum resolved shear stress (mRSS), stress triaxiality (η), and lode angle parameter ($\barθ$) values were tracked for ~300 grains under two different loading conditions: (1) Torsion-dominated loading (low NP) and (2) Tension-torsion loading (high NP) in equiatomic NiCoCr, a representative multicomponent face-centered cubic (FCC) superalloy. Overall, significant stress localizations existed within both samples as evidenced by the radial dependence of grain-resolved σ θZ , σ$^{grain}_{VM}$, and J 2 ; by comparison, I 1 , J 3 , η, and $\barθ$ metrics did not show discernible trends within the volume. These stress localizations reveal a complex interplay between axial and shear stress components (e.g., stress coupling) resulting in grain yielding near the sample surface largely driven by shear stress, whereas internal grain yielding was largely accommodated by axial stress. Grain-resolved stress localization trends were described well by the CPFE model, although some discrepancies in magnitude occurred, particularly for volumetric stress metrics (I 1 and η) due to initial type II residual stress distributions. The superposition of initial residual stress states onto CPFE grain-resolved data significantly improved model accuracy for η. This suggests that residual stresses more strongly influence the simulation of volumetric rather than deviatoric (yield) stress metrics.

36 MATERIALS SCIENCE↗

Vertical column dual-comb spectroscopy to a TBS

Open-path dual-frequency-comb spectroscopy (DCS) is a broadband, high spectral resolution, and high precision method for measuring gas concentrations over kilometer-scale paths. It has been used for detection and quantification of emissions of pollutants, hazardous gasses, and greenhouse gasses (GHGs). DCS has been shown to measure trace-gas mixing ratios with 0.14%-0.4% agreement between instruments. To achieve high signal-to-noise ratios (SNRs) with DCS, comb light is targeted onto a retroreflector at the end of the measurement path, which returns the signal to a detector co-located with the launch signal. Installing the retroreflector on a mobile platform such as a balloon or unmanned aerial vehicle (UAV) extends the capabilities of DCS by enabling variable path lengths, greater mobility, and access to higher altitudes. Mobile-target DCS has many uses in plume and leak detection, emissions modeling, and planetary boundary layer (PBL) studies. It is a promising method for observing vertical distributions of GHGs and mixing processes in the PBL, which are difficult to measure but important for pollution and climate monitoring as well as for understanding transport of gasses through the atmosphere. Tethered balloons are an intriguing platform because they enable longer flight durations and higher altitudes than easily obtainable with a UAV and thus allow for column measurements up to and above the PBL. New measurements completed in October/November 2024 reach the highest altitudes above ground level yet achieved by mobile-target DCS. For these measurements, the retroreflector was mounted on a 7 m diameter tethered helium balloon. An actively tracking gimbal on the ground holds the DCS launch telescope and keeps the 5 cm beam pointed onto the retroreflector while the balloon is lifted, lowered, and moved by wind and turbulence.

atmosphere↗

Automated Identification of Characteristic Droplet Size Distributions in Stratocumulus Clouds Utilizing a Data Clustering Algorithm

Abstract Droplet-level interactions in clouds are often parameterized by a modified gamma fitted to a “global” droplet size distribution. Do “local” droplet size distributions of relevance to microphysical processes look like these average distributions? This paper describes an algorithm to search and classify characteristic size distributions within a cloud. The approach combines hypothesis testing, specifically, the Kolmogorov–Smirnov (KS) test, and a widely used class of machine learning algorithms for identifying clusters of samples with similar properties: density-based spatial clustering of applications with noise (DBSCAN) is used as the specific example for illustration. The two-sample KS test does not presume any specific distribution, is parameter free, and avoids biases from binning. Importantly, the number of clusters is not an input parameter of the DBSCAN-type algorithms but is independently determined in an unsupervised fashion. As implemented, it works on an abstract space from the KS test results, and hence spatial correlation is not required for a cluster. The method is explored using data obtained from the Holographic Detector for Clouds (HOLODEC) deployed during the Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) field campaign. The algorithm identifies evidence of the existence of clusters of nearly identical local size distributions. It is found that cloud segments have as few as one and as many as seven characteristic size distributions. To validate the algorithm’s robustness, it is tested on a synthetic dataset and successfully identifies the predefined distributions at plausible noise levels. The algorithm is general and is expected to be useful in other applications, such as remote sensing of cloud and rain properties. Significance Statement A typical cloud can have billions of drops spread over tens or hundreds of kilometers in space. Keeping track of the sizes, positions, and interactions of all of these droplets is impractical, and, as such, information about the relative abundance of large and small drops is typically quantified with a “size distribution.” Droplets in a cloud interact locally, however, so this work is motivated by the question of whether the cloud droplet size distribution is different in different parts of a cloud. A new method, based on hypothesis testing and machine learning, determines how many different size distributions are contained in a given cloud. This is important because the size distribution describes processes such as cloud droplet growth and light transmission through clouds.

54 ENVIRONMENTAL SCIENCES↗

Global Methane Budget 2000–2020

Abstract. Understanding and quantifying the global methane (CH4) budget is important for assessing realistic pathways to mitigate climate change. CH4 is the second most important human-influenced greenhouse gas in terms of climate forcing after carbon dioxide (CO2), and both emissions and atmospheric concentrations of CH4 have continued to increase since 2007 after a temporary pause. The relative importance of CH4 emissions compared to those of CO2 for temperature change is related to its shorter atmospheric lifetime, stronger radiative effect, and acceleration in atmospheric growth rate over the past decade, the causes of which are still debated. Two major challenges in quantifying the factors responsible for the observed atmospheric growth rate arise from diverse, geographically overlapping CH4 sources and from the uncertain magnitude and temporal change in the destruction of CH4 by short-lived and highly variable hydroxyl radicals (OH). To address these challenges, we have established a consortium of multidisciplinary scientists under the umbrella of the Global Carbon Project to improve, synthesise, and update the global CH4 budget regularly and to stimulate new research on the methane cycle. Following Saunois et al. (2016, 2020), we present here the third version of the living review paper dedicated to the decadal CH4 budget, integrating results of top-down CH4 emission estimates (based on in situ and Greenhouse Gases Observing SATellite (GOSAT) atmospheric observations and an ensemble of atmospheric inverse-model results) and bottom-up estimates (based on process-based models for estimating land surface emissions and atmospheric chemistry, inventories of anthropogenic emissions, and data-driven extrapolations). We present a budget for the most recent 2010–2019 calendar decade (the latest period for which full data sets are available), for the previous decade of 2000–2009 and for the year 2020. The revision of the bottom-up budget in this 2025 edition benefits from important progress in estimating inland freshwater emissions, with better counting of emissions from lakes and ponds, reservoirs, and streams and rivers. This budget also reduces double counting across freshwater and wetland emissions and, for the first time, includes an estimate of the potential double counting that may exist (average of 23 Tg CH4 yr−1). Bottom-up approaches show that the combined wetland and inland freshwater emissions average 248 [159–369] Tg CH4 yr−1 for the 2010–2019 decade. Natural fluxes are perturbed by human activities through climate, eutrophication, and land use. In this budget, we also estimate, for the first time, this anthropogenic component contributing to wetland and inland freshwater emissions. Newly available gridded products also allowed us to derive an almost complete latitudinal and regional budget based on bottom-up approaches. For the 2010–2019 decade, global CH4 emissions are estimated by atmospheric inversions (top-down) to be 575 Tg CH4 yr−1 (range 553–586, corresponding to the minimum and maximum estimates of the model ensemble). Of this amount, 369 Tg CH4 yr−1 or ∼ 65 % is attributed to direct anthropogenic sources in the fossil, agriculture, and waste and anthropogenic biomass burning (range 350–391 Tg CH4 yr−1 or 63 %–68 %). For the 2000–2009 period, the atmospheric inversions give a slightly lower total emission than for 2010–2019, by 32 Tg CH4 yr−1 (range 9–40). The 2020 emission rate is the highest of the period and reaches 608 Tg CH4 yr−1 (range 581–627), which is 12 % higher than the average emissions in the 2000s. Since 2012, global direct anthropogenic CH4 emission trends have been tracking scenarios that assume no or minimal climate mitigation policies proposed by the Intergovernmental Panel on Climate Change (shared socio-economic pathways SSP5 and SSP3). Bottom-up methods suggest 16 % (94 Tg CH4 yr−1) larger global emissions (669 Tg CH4 yr−1, range 512–849) than top-down inversion methods for the 2010–2019 period. The discrepancy between the bottom-up and the top-down budgets has been greatly reduced compared to the previous differences (167 and 156 Tg CH4 yr−1 in Saunois et al. (2016, 2020) respectively), and for the first time uncertainties in bottom-up and top-down budgets overlap. Although differences have been reduced between inversions and bottom-up, the most important source of uncertainty in the global CH4 budget is still attributable to natural emissions, especially those from wetlands and inland freshwaters. The tropospheric loss of methane, as the main contributor to methane lifetime, has been estimated at 563 [510–663] Tg CH4 yr−1 based on chemistry–climate models. These values are slightly larger than for 2000–2009 due to the impact of the rise in atmospheric methane and remaining large uncertainty (∼ 25 %). The total sink of CH4 is estimated at 633 [507–796] Tg CH4 yr−1 by the bottom-up approaches and at 554 [550–567] Tg CH4 yr−1 by top-down approaches. However, most of the top-down models use the same OH distribution, which introduces less uncertainty to the global budget than is likely justified. For 2010–2019, agriculture and waste contributed an estimated 228 [213–242] Tg CH4 yr−1 in the top-down budget and 211 [195–231] Tg CH4 yr−1 in the bottom-up budget. Fossil fuel emissions contributed 115 [100–124] Tg CH4 yr−1 in the top-down budget and 120 [117–125] Tg CH4 yr−1 in the bottom-up budget. Biomass and biofuel burning contributed 27 [26–27] Tg CH4 yr−1 in the top-down budget and 28 [21–39] Tg CH4 yr−1 in the bottom-up budget. We identify five major priorities for improving the CH4 budget: (i) producing a global, high-resolution map of water-saturated soils and inundated areas emitting CH4 based on a robust classification of different types of emitting ecosystems; (ii) further development of process-based models for inland-water emissions; (iii) intensification of CH4 observations at local (e.g. FLUXNET-CH4 measurements, urban-scale monitoring, satellite imagery with pointing capabilities) to regional scales (surface networks and global remote sensing measurements from satellites) to constrain both bottom-up models and atmospheric inversions; (iv) improvements of transport models and the representation of photochemical sinks in top-down inversions; and (v) integration of 3D variational inversion systems using isotopic and/or co-emitted species such as ethane as well as information in the bottom-up inventories on anthropogenic super-emitters detected by remote sensing (mainly oil and gas sector but also coal, agriculture, and landfills) to improve source partitioning. The data presented here can be downloaded from https://doi.org/10.18160/GKQ9-2RHT (Martinez et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

Technical Performance and Cost Optimization of Unobtrusive Multi-static Serial LiDAR Imager (UMSLI) for Wide-area Surveillance and Identification of Marine Life at Marine Energy Installations

Florida Atlantic University developed an underwater optical monitoring system prototype - Unobtrusive Multi-static Serial LiDAR Imager (UMSLI) - suitable for marine energy full project lifecycle observation (baseline, commissioning, and decommissioning), with an automated real-time classification of marine animals. With precursor 2014 DOE funding (award DE-EE0006787), a prototype UMSLI was demonstrated in a controlled laboratory environment and achieved a TRL 6. This EERE DE-EE0007828 award aimed to both increase the TRL of the UMSLI by improving the technology performance (e.g., increase distance of marine animal target detection capability, add additional species classification capabilities, improve the system performance during the day, etc.) and reduce the cost. The UMSLI presents a novel application of underwater distributed Light Detection And Ranging (LiDAR), an advanced remote sensing method that uses light in the form of laser pulses for an application, this paired with an algorithm provides 360 degrees underwater detection, imaging, and classification of marine life. This solution for underwater monitoring of biota preserves the advantages of traditional optical and acoustic solutions while overcoming many associated disadvantages for marine energy site environmental monitoring, such as difficulties in species detection and classification in low light or night and turbid environments. This new approach is a purposefully designed, reconfigurable adaptation of an existing class 3B laser technology into one UMSLI instrument that can be easily mounted on or around different classes of marine energy equipment, such as devices to capture ocean current energy or wave energy. The system uses relatively low average power, utilizes far-red (> 635nm) laser illumination to be invisible and eye-safe to marine animals, is compact, and cost-effective. The equipment is designed for long-term, maintenance-free operations (i.e., current design is targeting more than 7 days continuous operation), to inherently generate a sparse primary dataset that only includes detected anomalies, and to allow robust real-time automated animal classification and identification with a low data bandwidth requirement. The technology’s overarching goal for application, is a system that can be deployed to collect pre-installation baseline species observations at a proposed marine energy deployment site with minimal post-processing overhead. The envisioned system will also produce high-resolution imagery of marine animals through a wide range of conditions and support automated tracking and notification of the presence of managed animals within established perimeters of marine energy equipment to satisfy deployed marine energy projects’ endangered and threatened species monitoring requirements. Through the current project, we demonstrated the UMSLI prototype in an operational environment and increased the UMSLI’s Technology Readiness Level from 6 to 7. The project resulted in many novel technologies, including: 1) an eye-safe red laser based, low-cost LiDAR system that can detect targets up to 10 meters distance; 2) the GAN-based machine learning underwater LiDAR image enhancement technique (this is the first known application of GAN technique in underwater LiDAR); 3) LiDAR-based real-time automated detection capabilities; and 4) a template matching based automated classification tool. These technologies build a solid foundation for future efforts to develop an extended range electro-optical monitoring system suitable for marine energy deployments. In addition to technology development, a key lesson learned is that addressing regulatory and safety requirements must be front and center in any marine energy monitoring applications.

16 TIDAL AND WAVE POWER↗

Atmospheric H 2 observations from the NOAA Cooperative Global Air Sampling Network

Abstract. The NOAA Global Monitoring Laboratory (GML) measures atmospheric hydrogen (H2) in grab samples collected weekly as flask pairs at over 50 sites in the Cooperative Global Air Sampling Network. Measurements representative of background air sampling show higher H2 in recent years at all latitudes. The marine boundary layer (MBL) global mean H2 was 552.8 ppb in 2021, 20.2 ± 0.2 ppb higher compared to 2010. A 10 ppb or more increase over the 2010–2021 average annual cycle was detected in 2016 for MBL zonal means in the tropics and in the Southern Hemisphere. Carbon monoxide measurements in the same-air samples suggest large biomass burning events in different regions likely contributed to the observed interannual variability at different latitudes. The NOAA H2 measurements from 2009 to 2021 are now based on the World Meteorological Organization Global Atmospheric Watch (WMO GAW) H2 mole fraction calibration scale, developed and maintained by the Max Planck Institute for Biogeochemistry (MPI-BGC), Jena, Germany. GML maintains eight H2 primary calibration standards to propagate the WMO scale. These are gravimetric hydrogen-in-air mixtures in electropolished stainless steel cylinders (Essex Industries, St. Louis, MO), which are stable for H2. These mixtures were calibrated at the MPI-BGC, the WMO Central Calibration Laboratory (CCL) for H2, in late 2020 and span the range 250–700 ppb. We have used the CCL assignments to propagate the WMO H2 calibration scale to NOAA air measurements performed using gas chromatography and helium pulse discharge detector instruments since 2009. To propagate the scale, NOAA uses a hierarchy of secondary and tertiary standards, which consist of high-pressure whole-air mixtures in aluminum cylinders, calibrated against the primary and secondary standards, respectively. Hydrogen at the parts per billion level has a tendency to increase in aluminum cylinders over time. We fit the calibration histories of these standards with zero-, first-, or second-order polynomial functions of time and use the time-dependent mole fraction assignments on the WMO scale to reprocess all tank air and flask air H2 measurement records. The robustness of the scale propagation over multiple years is evaluated with the regular analysis of target air cylinders and with long-term same-air measurement comparison efforts with WMO GAW partner laboratories. Long-term calibrated, globally distributed, and freely accessible measurements of H2 and other gases and isotopes continue to be essential to track and interpret regional and global changes in the atmosphere composition. The adoption of the WMO H2 calibration scale and subsequent reprocessing of NOAA atmospheric data constitute a significant improvement in the NOAA H2 measurement records.

Pétron, Gabrielle↗

Investigation of Best-Practices and Computationally Inexpensive Radiative Exchange Models for Discrete Element Method Modeling of Aluminosilicate Particles in Concentrating Solar Power Environments

Chemically inert, aluminosilicate based particles have been investigated as both a thermal transport and sensible energy storage medium for concentrating solar power facilities. These particles will experience a wide range of operating temperatures (300-1000 K) and handling conditions (dense to dilute falling particle curtains, dense granular flows, or dense structures), requiring specially-designed and optimized infrastructures. The relative influence of collisional and frictional interactions between particles varies based on temperature-dependent particulate properties and greatly impacts the bulk, granular flow behavior. These underlying physics are captured using discrete element method modeling tools. However, this modeling method is computationally expensive as each particle position and interaction is tracked during the simulation. These modeling methods are further complicated by introducing temperature-dependent particle properties, high-temperature radiative exchange, and directional irradiation sources experienced by granular flows in concentrating solar power environments. In this study, coupled experimental and numerical slump testing of aluminosilicate particles was performed and computationally efficient radiative exchange models were evaluated to establish best-practices for discrete element method models for concentrating solar power environments. The three particle types investigated included Carbobead HSP 30 /60, Carbobead CP 30/60, and Granusil 4030. Existing modeling limitations and computationally-efficient multi-modal heat transfer models were evaluated using Aspherix®, a commercial discrete element method software. High-temperature (< 1073 K) slump testing of aluminosilicate particles was performed to investigate the deviation between experimentally-observed and numerically-predicted angles of repose introduced by computation-time reduction practices including the relaxation of the particle elastic modulus and coarse-graining. Coarse-graining is used to use a single modeled particle that is representative of a collection of smaller particles, decreasing the computational cost at the expense of geometric accuracy. Additionally, relaxation of the elastic modulus is used to reduce computational time at the expense of an increased, modeled particle overlap. Prior studies have determined that aluminosilicate particles retain a high elastic modulus at high temperatures (< 1073 K), requiring small simulation timesteps to ensure resolved contact forces resemble appropriate solid mechanics. A parametric study was performed to evaluate the influence of computation time improvements on the deviation between experimental and modeled angle of repose across high temperatures < 1073 K. Additionally, numerical case studies were performed on candidate particle systems at varying porosities and temperatures. These studies were performed to investigate the influence of computationally-efficient radiative-exchange modeling methods coupled to Aspherix® on modeled accuracy and computation time. The recently-developed distance-based approximation was evaluated in estimating radiative exchange between particles and participating surfaces located in close proximity. The distance based approximation was developed to use tabulated estimates of the radiative distribution factor between individual particles and surfaces in close proximity (< 40 particle radii). These methods were expanded to the aluminosilicate particles of interest, including the influence of particle size distributions. To capture radiative exchange between particles and surfaces not in close proximity (> 40 particle radii) and to capture the absorption of directional irradiation from concentrating solar resources, a volumetrically-averaged radiative distribution factor was calculated between the modeled granular flow and surfaces using Monte Carlo ray-tracing for participating media. Volume-averaged absorption and scattering coefficients were predicted using a volumetric discretization of the modeled domain with monodisperse approximations based on geometric optics and experimentally-determined scattering phase functions for aluminosilicate particles.

14 SOLAR ENERGY↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗