Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data reconciliation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

The utility of Bayesian data reconciliation for separations

Data reconciliation methods for separation processes typically rely on classical statistical approaches to generate estimates of true mass flow rates from measurements. Knowledge regarding the uncertainty of these estimates has value in decision making, but is often not acquired. Bayesian approaches intrinsically quantify uncertainty; however, literature for Bayesian data reconciliation of separation processes is scarce. This publication outlines two Bayesian data reconciliation models and provides details for how the models were implemented for the BayesMassBal (V 1.0.0) software package written in R. To demonstrate the advantages of this approach for data reconciliation, the models were first applied to simulated data and then compared to a classical model through a Monte Carlo experiment. In this example, the Bayesian models were found to provide more accurate estimates of the simulated data, while also providing quantitative information on the estimate uncertainty. To demonstrate the use of the technique in a practical problem, the models were also applied to real data collected from a pilot-scale rare earth solvent extraction process. Here, this publication provides a small window into how Bayesian methods can be used for data reconciliation, but findings suggest Bayesian data reconciliation models for separation processes have distinct advantages over classical alternatives.

01 COAL, LIGNITE, AND PEAT↗

Dynamic Modeling, Parameter Estimation, and Data Reconciliation of a Supercritical Pulverized Coal-Fired Boiler

A high-fidelity, spatially distributed dynamic model of a supercritical pulverized coal-fired boiler was developed for calculating transient thermal profiles of the boiler tubes, which are often not measured or unmeasurable in industrial settings due to the harsh operating conditions. In this work, rigorous models of the subcritical and supercritical steam properties, which can be highly nonlinear, were used to capture the transition between two-phase and one-phase flow, respectively, as supercritical boilers frequently transition through the critical point during load-following operation. The transient boiler model was also used for data reconciliation and parameter estimation for the purpose of model validation against industrial data. An approach was developed to reduce the size of large-scale reconciliation problem by focusing on the implementation of bias terms for the reconciled variables, which can be estimated in place of the reconciled variables themselves. In addition, a functional approximation of the reconciliation problem was defined so that it can be more simply optimized. Combining this functional approximation optimization with reformulation of the reconciliation problem considering measurement biases can significantly improve the tractability of many such large-scale problems. The validated boiler model was then used to study transient responses under load-following operation with specific attention on the calculation of unmeasurable variables.

42 ENGINEERING↗

Microwave-Assisted Plastic Upcycling: Dynamic Data Reconciliation, Parameter Estimation, and Kinetic Modeling

Microwave (MW)-assisted catalytic pyrolysis offers a promising pathway for efficient plastic upcycling. This work develops an integrated modeling framework combining dynamic data reconciliation, a temperature-dependent rate model, and a yield model to represent the time-varying production rate of components in MW-assisted LDPE pyrolysis conducted in a batch reactor. An Arrhenius-type rate model with a temperature-dependent reaction order is developed. A biexponential correlation is proposed for the yield of gaseous products that enables to capture the evolving product formation behavior during conversion. In the yield correlation, one term is used to represent the initial increase in yield, reflecting the rapid formation of intermediate or primary products at the early stages of the reaction when a larger fraction of the reactant remains available. As conversion progresses, the influence of this term gradually diminishes. The other term accounts for the subsequent decrease in the predicted yield, representing secondary reactions such as further cracking or coke formation that reduce the concentration of certain products at higher conversion. The model is found to accurately represent reconciled experimental flow rate profiles from an in-house MW-assisted catalytic batch reactor for major products, including ethylene, ethane, 1-butene, and benzene, across 250−350 °C. Ethylene remains the dominant product but decreases from about 41.95% at 250 °C to 30.14% at 350 °C, while heavier products increase significantly, with 1-butene rising to nearly 8.37% and benzene reaching 2.17% at intermediate temperatures. The model shows that the ethylene production rate can be maximized at around 270 °C. The models developed in this work can be utilized for process optimization, reactor design and scale-up of microwave-assisted plastic conversion technologies, and economic analysis.

Damahe, Harish [West Virginia Univ., Morgantown, W↗

Santa Barbara Desalination Digital Twin Technical Report

A digital twin is a computer model that has sufficient breadth and fidelity to enable diagnostic and predictive analyses for outputs of interest. We have constructed a digital twin of the SB desal plant using the ProteusLib software developed for NAWI under task 4.2.2, and specifically the ProteusLib models of pumps, filters, mixers, splitters, and reverse osmosis units and a seawater property model. The steps needed to configure this digital twin with the real conditions of the SB desal plant are: (1) Data reconciliation, which is gathering the data and plant connectivity and ensure that all relevant values are present and in the correct units and validating the data using material and energy balances, and correct for errors, gaps, or deviations; (2) Parameter estimation, which is matching model parameters with data from previous step, filling missing parameters with initial estimates, then iteratively refining estimated values. Once we completed the digital twin, we used the optimization capabilities of ProteusLib to compute outputs of interest over a chosen range of inputs. The primary output of interest to the owners of the plant is its specific energy consumption (kWh/acre-foot). The remainder of this report will provide more detail on the data reconciliation, parameter estimation, and analyses we performed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Heat Loss Characteristics and Energy Use of Piperazine with the Advanced Stripper (PZAS) at the UT-SRP Pilot Plant

Heat duty and heat loss were measured at the pilot plant at UT Austin. Heat loss was measured with energy balances using water. Heat loss was studied using surface temperature measurements over 68 different locations at the pilot plant. Surface temperature measurements indicated that bare metal surfaces were the primary source of heat loss from the pilot plant. Heat loss from bare metal surfaces was at least 50% controlled by natural convection and in many cases was as high as 75% natural convection controlled. Measured heat loss at the pilot plant was 20100 BTU/hr and overall heat loss did not show any dependence on the heat rate of the plant. Compared to PZAS™ at the National Carbon Capture Center, heat loss relative to heat rate was higher at 38%. The relative heat loss at the pilot plants was found to decrease by 20% per MW of added capacity. Measured net heat duty was found to be dependent on measured lean loading and cold rich bypass flow rate. Measured net heat duty was between 2.2 and 2.4 GJ/tonne at an optimum lean loading of 0.2–0.21 mol/mol. Data reconciliation by Aspen Plus® Data Fit™ underpredicted CO2 flow rate by 20% due to an overprediction of lean loading by 19%, indicating a necessary change in thermodynamic parameters in the model. This resulted in an over prediction of heat duty by 33% on average.

Amine scrubbing, stripper, energy requirement, hea↗

Design of Experiments for Dynamic Test Runs in Solvent-Based CO 2 Capture Pilot Plants

Test runs in the pilot plants consume significant resources, and therefore, the learning from test runs should be maximized. Test runs conducted in the pilot plants are often steady state. It takes several hours for reaching steady-state in the pilot plants, and thus, the duration of the test runs needs to be long even for collecting few steady-state data points. On the other hand, a large number of measurements can be collected through dynamic test runs in a short span of time. This paper presents a systematic design of dynamic experiments (DoDEs) for identifiability of model parameters, which is achieved by persistently exciting the inputs signals. A pseudorandom binary sequence (PRBS) is designed as the input signal for DoDE due to its efficiency in obtaining sufficient spectral content. However, due to the long sequence size of the PRBS signal, a Schroeder-phase input signal, which is a multisine signal, is also designed. Tests for both types of signals are run in the Pilot Solvent Test Unit (PSTU) at the National Carbon Capture Center in Wilsonville, Alabama. The transient data are used to solve dynamic data reconciliation and parameter estimation problem. The estimated parameters are found to be not only superior to those estimated from using data collected from hundreds of steady-state test runs in a nonreactive (air–water) system, but the parameters could be estimated by using the dynamic data collected for about 24 h from the pilot plant for the MEA-H 2 O–CO 2 system.

CO2 capture↗

Leveraging Radiofrequency Identification Success Beyond Hazardous Material Inventory Management at a National Laboratory

Effective inventory management can be overshadowed by conflicting priorities in organizational procedures, particularly in research-focused institutions such as national laboratories that handle expensive, delicate, and hazardous materials. Here, this study investigated the potential of radiofrequency identification (RFID) technology, currently used for hazardous chemical inventory, in applications with higher metal interference and absorption, specifically pressure release device (PRD) compliance and nuclear container management, at Lawrence Livermore National Laboratory (LLNL). This study was done to document best practices to enhance inventory identification speeds for inventory reconciliation and inventory recall and to explore optimal configurations for RFID implementation compared to traditional manual methods of equipment management. Tests were conducted to determine the ideal RFID tag orientation (read at angles of 0°, 90°, and 270°), various container layouts (linear, separated, curved, operational), and ID methods such as manual, barcode, and RFID performing three trials per method per orientation. Results indicated that 0° was the optimal read angle for minimizing metallic interference, and the operational and curved arrangements significantly outperformed the linear and separated configurations in read speed. 3D printed mounts were developed and tested, increasing the read range of the RFID reader by up to 235% in cases of high metallic interference. The RFID technology demonstrated an average speed increase of 65% over a simplified manual identification, which supports the conclusion that RFID is a more efficient method for large hazardous inventory management and equipment reconciliation. Additionally, capturing meta-data, such as location and date, can be used to query for inventory recall and automated updating of record information.

42 ENGINEERING↗

Coordinated resource allocation to plant growth–defense tradeoffs

Summary Plant resource allocation patterns often reveal tradeoffs that favor growth (G) over defense (D), or vice versa. Ecologists most often explain G–D tradeoffs through principles of economic optimality, in which negative trait correlations are attributed to the reconciliation of fitness costs. Recently, researchers in molecular biology have developed ‘big data’ resources including multi‐omic (e.g. transcriptomic, proteomic and metabolomic) studies that describe the cellular processes controlling gene expression in model species. In this synthesis, we bridge ecological theory with discoveries in multi‐omics biology to better understand how selection has shaped the mechanisms of G–D tradeoffs. Multi‐omic studies reveal strategically coordinated patterns in resource allocation that are enabled by phytohormone crosstalk and transcriptional signal cascades. Coordinated resource allocation justifies the framework of optimality theory, while providing mechanistic insight into the feedbacks and control hubs that calibrate G–D tradeoff commitments. We use the existing literature to describe the coordinated resource allocation hypothesis (CoRAH) that accounts for balanced cellular controls during the expression of G–D tradeoffs, while sustaining stored resource pools to buffer the impacts of future stresses. The integrative mechanisms of the CoRAH unify the supply‐ and demand‐side perspectives of previous G–D tradeoff theories.

Monson, Russell K.↗

Assessing the Potential Impact of Fugitive Methane Emissions on Offshore Platform Safety

One of the biggest risks to safety on offshore platform safety is the ignition of high-pressure natural gas streams. Currently, the size and number of fugitive emissions on offshore platforms is unknown and methods used to detect fugitives have significant shortcomings. To investigate the frequency, size, and potential impact of fugitives, a data collection exercise was conducted using incidents reported, leak survey data, and independent measurements. The size and number of fugitives on offshore facilities were simulated to investigate likely areas of safety concern. Incident reports indicate in 2021 there were 113 reports of gas leaks on 1119 offshore facilities, suggesting 0.02 fugitives per Type 1 facility (older, shallow-water platforms) and 0.31 fugitives per Type 2 facility (larger deeper-water facilities). Leak survey data report 12 fugitives per Type 1 facility (average emission 0.6 kg CH 4 h −1 leak −1 ) and 15 fugitives per Type 2 facility (average emission 1.5 kg CH 4 h −1 leak −1 ). Reconciliation of direct measurements with a bottom-up model suggests that the number of fugitive emissions generated from the leak report data is an underestimate for Type 1 platforms (44 fugitives facility −1 ; average emission 0.6 kg CH 4 h −1 leak −1 ) and in general agreement for the Type 2 platforms (15 fugitives facility −1 ; average emission 1.5 kg CH 4 h −1 leak −1 ). Analysis of the fugitive emission rates on an offshore platform suggests that gas will not collect to explosive concentration if any air movement is present (>0.36 mph); however, large volumes of air (~600 m 3 ) near representative leaks on the working deck could become explosive in hour-long zero-wind conditions. We suggest that wearable technology could be employed to indicate gas build up, safety regulations amended to consider low-wind conditions and real-world experiments are conducted to test assumptions of air mixing on the working deck.

explosion↗

DOE CESER 6 GHz Interference Study

DOE CESER has sponsored Idaho National Laboratory (INL) to conduct an objective and independent study of potential 6 GHz interference from outdoor operation of unlicensed devices in the 6 GHz band on fixed service (FS) microwave communication links operated by electrical sector incumbents in that band. INL is collaborating with University of Notre Dame (UND), Electric Power Research Institute (EPRI), Lockard & White, Southern Company, and AT&T, to gather data with real-world 6 GHz interference experiments and identify (1) the potential for interference from unlicensed devices and (2) the interference necessary to cause harm to the incumbents. In addition to the functional assessment, a security assessment of the FCC mandated Automatic Frequency Coordination (AFC) System to regulate use of unlicensed 6 GHz standard power devices is also being conducted. A major objective is to create a science-based and defensible methodology used to produce the necessary data and to derive objective conclusions. This proven methodology can then be used to produce objective data and conclusions for other spectrum bands with similar incumbent uses including 4.4 – 4.9 GHz and 7.125 – 7.4 GHz, identified in the reconciliation bill that was adopted on July 4, 2025, as well as the National Spectrum Strategy discussions that are ongoing. This report contains 6 GHz field experiments and findings in the following real-world scenarios with commercial unlicensed standard power (SP) 6 GHz devices regulated by Automated Frequency Coordination (AFC): • University of Notre Dame (UND) Stadium with a capacity of 80,000 spectators, where Wi-Fi operating in 6 GHz has been deployed recently • Southern Company 6 GHz FS microwave link between Columbus and Fortson, Georgia Following are the following key findings from this study. 1. The AFC is under-protective of FS when line-of-sight exists along the path centerline. Data collected at Southern’s 6 GHz fixed link site shows significant erosion of as much as 21.4-24.4 dB of under-protection that can lead to potential service degradation under typical operating conditions. This first key finding is most likely the result of erroneous use of the RF propagation model. INL will collaborate with EPRI and the AFC Functional Requirements Working Group to submit a change request to the WinnForum TS-1014 standard towards correct use of the propagation model by the AFC. 2. There is additive interference effect of about 3 dB from nearly equal power interferers measured from simultaneous operation of two SP AP's operating co-channel with the FS receive from different locations along the path. This second key finding should be used to add the impact of additive interference of operation of multiple APs in the same geographical area, to the next generation of AFCs. INL will collaborate with FCC on the need for the AFC to consider additive interference. We also recommend that additional experiments are conducted on 6 GHz spectrum interference to further improve the AFC operation as the number of outdoor Wi-Fi devices continues to increase. These proposed steps and recommendations will make the co-existence of the incumbents and the 6 GHz outdoor Wi-Fi providers possible without any impact on the incumbents with a win-win outcome for all.

6 GHz↗

Evaluation of New Additions to OLI Software in Predicting Mercuric and Mercurous Species in Liquid Waste Operations

Speciation of mercury during the pretreatment steps of tank waste processing is critical to successful mercury removal prior to vitrification during Liquid Waste Operations (LWO) at SRS. OLI software has been used to predict mercury speciation and activity throughout LWO. The OLI software operates based on a thermodynamic framework called the Mixed Solvent Electrolyte (MSE) framework. The MSE framework allows prediction in theoretically infinitely dilute to concentrated mixtures (e.g., purely solute solutions). Before modification to the MSE framework databanks, certain critical mercury species were missing in the MSE databank, and some thermodynamic data needed to be updated for the OLI software to accurately predict mercury chemical species in SRS waste tanks. To better reflect streams across LWO, new mercury species were integrated into the MSE database. To evaluate the changes to the OLI MSE framework per the Technical Task Request (TTR) and the Task Technical and Quality Assurance Plan (TTQAP), waste stream compositions from Tanks 38, 43, and Tank 50 decontaminated salt solution (DSS) were used as model inputs. Models were developed and executed using both the old and new databases. Compositional analyses from caustic Tank 50 DSS and caustic Tanks 38 and 43 were used as the input streams. These streams represent the most comprehensive chemical data sets where both mercury and tank constituents were measured together. Results for Tank 50 DSS predict HgO as the predominant species in both databases. Both methyl and dimethyl Hg species are present when the new database is ‘on’ and are not predicted with the new database turned ‘off’. The new database predicts a greater amount of HgO and a greater fraction of it in the solid phase. Pourbaix diagrams (potential vs. pH) generated for each Tank 50 DSS were identical regardless of which database was used. Elemental Hg and HgO were predicted in the water stable region under basic conditions. Tanks 38 and 43 follow similar trends as the Tank 50 DSS models. Unlike Tanks 38 and 50 DSS, the Tank 43 Pourbaix plot shows a region of stability for an aqueous HgOHCO3 - species between approximately pH 7-11. In all streams, when MeHg+ is included in the inputs, the new database predicts aqueous MeHgOH as the dominant species. If elemental or dimethyl mercury is in the waste stream, the new database model predicts they are unchanged and remain in those states and quantities. Additionally, the total mercury values are reported for both the measured input data and the OLI output data for all considered tanks. The summary indicates that the percentage error between the measured and calculated values is less than 1% in all cases The reconciliations and generation of the Pourbaix diagrams for Tank 50 DSS took approximately ten times longer with the new database ‘on’. In addition, over the course of that time, models with the new database ‘on’ were more likely to crash or display an error. Some modest performance improvements were noted when modeling with an i7 processor versus an i5. An example error is found in Appendix A. Furthermore, Appendix B provides V&V for two chemical systems analyzed with the OLI software, results were satisfactory. It is recommended to utilize the new databases (i.e., HCO.ddb and SR-Hg.ddb) in future Savannah River Mission Completion applications of OLI to represent pseudo steady-state. Furthermore, the integration and utilization of the new databases (i.e., HCO.ddb and SR-Hg.ddb) in modeling applications (e.g., Aspen) is also recommended.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

An iterative bidirectional gradient boosting approach for CVR baseline estimation

Here this paper presents a novel Iterative Bidirectional Gradient Boosting Model (IBi-GBM) for estimating the baseline of Conservation Voltage Reduction (CVR) programs. In contrast to many existing methods, we treat CVR baseline estimation as a missing data retrieval problem. The approach involves dividing the load and its corresponding temperature profiles into three periods: pre-CVR, CVR, and post-CVR. To restore the missing load profile during the CVR period, the method employs a three-step process. First, a forward-pass GBM is executed using data from the pre-CVR period as inputs. Subsequently, a backward-pass GBM is applied using data from the post-CVR period. The two restored load profiles are reconciled, considering pre-calculated weights derived from forecasting accuracy, and only the leftmost and rightmost points are retained. The newly restored points are then included as inputs for the subsequent iteration. This iterative procedure continues until the original load data in the CVR period is fully restored. We develop IBi-GBM using actual smart meter and Supervisory Control and Data Acquisition (SCADA) data. Our results demonstrate that IBi-GBM exhibits robust performance across various data resolutions and in different seasons and outperforms existing methods by achieving a 1-2% reduction in normalized Root Mean Square Error (nRMSE).

42 ENGINEERING↗

Progress in modeling hydrogen assisted ammonia oxidation with new experiments and a further reconciliation of the NH 3 + OH rate constant

Hydrogen-assisted oxidation of ammonia in a premixed, laminar flow tubular reactor under reducing conditions was investigated experimentally and through chemical kinetic modeling. Due to its impact on the competition for OH among ammonia and hydrogen, the rate constant for NH 3 + OH (R1) was determined through state-of-the-art theoretical kinetics calculations, employing composite energies that include the effects of higher order electronic excitations on the electronic energies along the variational reaction path and treating the limitations in the kinetics posed by the passage through a hydrogen-bonded complex. The resulting rate constant was in close agreement with the recent experimental value from Zaczek et al. (2025), settling a long-term dispute about the high-temperature value of k 1 and confirming within 20% the value previously used in modeling. The chemical kinetic model, with no other changes, captured well measured concentrations of NH 3 , H 2 , NO, and N 2 O from flow reactor oxidation of NH 3 /H 2 at slightly reducing conditions over a range of temperature (900-1350 K) and NH 3 /H 2 ratios (0.5-2.0). Comparison of the present results with reported data from a non-premixed setup indicates that for laminar flow tubular reactors, the reactor configuration may have implications for the observed H 2 consumption due to the possibility of preferential oxidation during mixing.

Ab initio theory↗

An independent analysis of bias sources and variability in wind plant pre-construction energy yield estimation methods

The wind resource assessment community has long had the goal of reducing the bias between wind plant pre-construction energy yield assessment (EYA) and the observed annual energy production (AEP). This comparison is typically made between the 50% probability of exceedance (P50) value of the EYA and the long-term corrected operational AEP (hereafter OA P50), and is known as the P50 bias. The industry has critically lacked an independent analysis of bias reduction investigated across multiple consultants to identify the greatest sources of uncertainty and variance in the EYA process and the best opportunities for uncertainty reduction. The present study addresses this gap by benchmarking consultant methodologies against each other and against operational data at a scale not seen before in industry collaborations. We consider data from 10 wind plants and evaluate discrepancies between eight consultancies in the steps taken from estimates of gross to net energy. Consultants tend to overestimate the gross energy produced at the turbines and then compensate by further overestimating downstream losses, leading to a mean P50 bias near zero, still with significant variability among the individual wind plants. Within our data sample, we find that consultant estimates of all loss categories, except environmental losses, tend to reduce the project-to-project variability of the P50 bias. The disagreement between consultants, however, remains flat throughout the addition of losses. Finally, we find that differences in consultants’ estimates of project performance can lead to differences up to $10/MWh in the levelized cost of energy for a wind plant.

Todd, Austin C.↗