Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Network parameter error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

Real-Time High-Accuracy Digital Wireless Time, Frequency, and Phase Calibration for Coherent Distributed Antenna Arrays

his work presents a fully-digital high-accuracy real-time calibration procedure for frequency and time alignment of open-loop wirelessly coordinated coherent distributed antenna array (CDA) modems, enabling radio frequency (RF) phase coherence of spatially separated commercial off-the-shelf (COTS) software-defined radios (SDRs) without cables or external references such as the global navigation satellite system (GNSS). Building on previous work using high-accuracy spectrally-sparse time of arrival (ToA) waveforms and a multistep ToA refinement process, a high-accuracy two-way time transfer (TWTT)-based time–frequency coordination approach is demonstrated. Due to the two-way nature of the high-accuracy TWTT approach, the time and frequency estimates are Doppler and multipath tolerant, so long as the channel is reciprocal over the synchronization epoch. This technique is experimentally verified using COTS SDRs in a lab environment in static and dynamic scenarios and with significant multipath scatterers. Time, frequency, and phase stability were evaluated by beamforming over coaxial cables to an oscilloscope which achieved time and phase precisions of ~60– 70 ps , with median coherent gains above 99% using optimized coordination parameters, and a beamforming frequency root-mean-square error (RMSE) of 3.73 ppb in a dynamic scenario. Finally, experiments were conducted to compare the performance of this technique with previous works using an analog continuous-wave two-tone (CWTT) frequency reference technique in both static and dynamic settings.

Clock synchronization↗

Accelerating massively parallel hemodynamic models of coarctation of the aorta using neural networks

Comorbidities such as anemia or hypertension and physiological factors related to exertion can influence a patient’s hemodynamics and increase the severity of many cardiovascular diseases. Observing and quantifying associations between these factors and hemodynamics can be difficult due to the multitude of co-existing conditions and blood flow parameters in real patient data. Machine learning-driven, physics-based simulations provide a means to understand how potentially correlated conditions may affect a particular patient. Here, we use a combination of machine learning and massively parallel computing to predict the effects of physiological factors on hemodynamics in patients with coarctation of the aorta. We first validated blood flow simulations against in vitro measurements in 3D-printed phantoms representing the patient’s vasculature. We then investigated the effects of varying the degree of stenosis, blood flow rate, and viscosity on two diagnostic metrics – pressure gradient across the stenosis (ΔP) and wall shear stress (WSS) - by performing the largest simulation study to date of coarctation of the aorta (over 70 million compute hours). Using machine learning models trained on data from the simulations and validated on two independent datasets, we developed a framework to identify the minimal training set required to build a predictive model on a per-patient basis. We then used this model to accurately predict ΔP (mean absolute error within 1.18 mmHg) and WSS (mean absolute error within 0.99 Pa) for patients with this disease.

59 BASIC BIOLOGICAL SCIENCES↗

Fast and accurate reduced-order modeling of a MOOSE-based additive manufacturing model with operator learning

One predominant challenge in additive manufacturing (AM) is to achieve specific material properties by manipulating manufacturing process parameters during the runtime. Such manipulation tends to increase the computational load imposed on existing simulation tools employed in AM. The goal of the present work is to construct a fast and accurate reduced-order model (ROM) for an AM model developed within the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, ultimately reducing the time/cost of AM control and optimization processes. Our adoption of the operator learning (OL) approach enabled us to learn a family of differential equations produced by altering process variables in the laser’s Gaussian point heat source. More specifically, we used the Fourier neural operator (FNO) and deep operator network (DeepONet) to develop ROMs for time-dependent responses. Furthermore, we benchmarked the performance of these OL methods against a conventional deep neural network (DNN)-based ROM. Ultimately, we found that OL methods offer comparable performance and, in terms of accuracy and generalizability, even outperform DNN at predicting scalar model responses. The DNN-based ROM afforded the fastest training time. Furthermore, all the ROMs were faster than the original MOOSE model yet still provided accurate predictions. FNO had a smaller mean prediction error than DeepONet, with a larger variance for time-dependent responses. Unlike DNN, both FNO and DeepONet were able to simulate time series data without the need for dimensionality reduction techniques. Finally, the present work can help facilitate the AM optimization process by enabling faster execution of simulation tools while still preserving evaluation accuracy.

36 MATERIALS SCIENCE↗

AutoSourceID-Classifier: Star-galaxy classification using a convolutional neural network with spatial information

Aims.Traditional star-galaxy classification techniques often rely on feature estimation from catalogs, a process susceptible to introducing inaccuracies, thereby potentially jeopardizing the classification’s reliability. Certain galaxies, especially those not manifesting as extended sources, can be misclassified when their shape parameters and flux solely drive the inference. We aim to create a robust and accurate classification network for identifying stars and galaxies directly from astronomical images. Methods.The AutoSourceID-Classifier (ASID-C) algorithm developed for this work uses 32x32 pixel single filter band source cutouts generated by the previously developed AutoSourceID-Light (ASID-L) code. By leveraging convolutional neural networks (CNN) and additional information about the source position within the full-field image, ASID-C aims to accurately classify all stars and galaxies within a survey. Subsequently, we employed a modified Platt scaling calibration for the output of the CNN, ensuring that the derived probabilities were effectively calibrated, delivering precise and reliable results. Results.We show that ASID-C, trained on MeerLICHT telescope images and using the Dark Energy Camera Legacy Survey (DECaLS) morphological classification, is a robust classifier and outperforms similar codes such as SourceExtractor. To facilitate a rigorous comparison, we also trained an eXtreme Gradient Boosting (XGBoost) model on tabular features extracted by SourceExtractor. While this XGBoost model approaches ASID-C in performance metrics, it does not offer the computational efficiency and reduced error propagation inherent in ASID-C’s direct image-based classification approach. ASID-C excels in low signal-to-noise ratio and crowded scenarios, potentially aiding in transient host identification and advancing deep-sky astronomy.

Astronomy & Astrophysics↗

Attention-based 3D – convolutional neural network model for mechanical property predictions using visible light images in metal additive manufacturing

Additive manufacturing (AM), while commonly used for rapid prototyping and creating components with complex geometries, has not been widely adopted for critical applications across the aerospace, automotive, defense, energy, and medical industries. This is, in part, due to the challenges of controlling flaws and uncertainty in the mechanical behavior of additively manufactured components. In recent years, there has been an increase in research aimed at predicting the final mechanical properties of additively manufactured components during the printing process. To address these issues, a 3D-CNN model was trained using low-cost in situ visible-light camera data, anomaly classifications, and the chosen process parameters to predict the ultimate tensile strength (UTS), yield strength (YS), total elongation (TE), and uniform elongation (UE). The 3D-CNN layers of the model employed attention mechanisms to prioritize features in the data, thereby improving prediction accuracy. Furthermore, the effect of each process parameter and anomaly class is investigated using attention-based dynamic sigmoid weighted gates to interpret the influence each class has on the final prediction. Different combinations of the in situ data were fed into the 3D-CNN, with varying amounts of image layers, to determine the ideal combination for predicting mechanical properties in situ. Here, the 3D-CNN model achieved mean absolute percentage errors (MAPE) below 5% for both UTS and YS while using only a single camera input and under half of the available image layers.

36 MATERIALS SCIENCE↗

Deep Learning Approaches to Surrogates for Solving the Diffusion Equation for Mechanistic Real-World Simulations

In many mechanistic medical, biological, physical, and engineered spatiotemporal dynamic models the numerical solution of partial differential equations (PDEs), especially for diffusion, fluid flow and mechanical relaxation, can make simulations impractically slow. Biological models of tissues and organs often require the simultaneous calculation of the spatial variation of concentration of dozens of diffusing chemical species. One clinical example where rapid calculation of a diffusing field is of use is the estimation of oxygen gradients in the retina, based on imaging of the retinal vasculature, to guide surgical interventions in diabetic retinopathy. Furthermore, the ability to predict blood perfusion and oxygenation may one day guide clinical interventions in diverse settings, i.e., from stent placement in treating heart disease to BOLD fMRI interpretation in evaluating cognitive function (Xie et al., 2019; Lee et al., 2020). Since the quasi-steady-state solutions required for fast-diffusing chemical species like oxygen are particularly computationally costly, we consider the use of a neural network to provide an approximate solution to the steady-state diffusion equation. Machine learning surrogates, neural networks trained to provide approximate solutions to such complicated numerical problems, can often provide speed-ups of several orders of magnitude compared to direct calculation. Surrogates of PDEs could enable use of larger and more detailed models than are possible with direct calculation and can make including such simulations in real-time or near-real time workflows practical. Creating a surrogate requires running the direct calculation tens of thousands of times to generate training data and then training the neural network, both of which are computationally expensive. Often the practical applications of such models require thousands to millions of replica simulations, for example for parameter identification and uncertainty quantification, each of which gains speed from surrogate use and rapidly recovers the up-front costs of surrogate generation. We use a Convolutional Neural Network to approximate the stationary solution to the diffusion equation in the case of two equal-diameter, circular, constant-value sources located at random positions in a two-dimensional square domain with absorbing boundary conditions. Such a configuration caricatures the chemical concentration field of a fast-diffusing species like oxygen in a tissue with two parallel blood vessels in a cross section perpendicular to the two blood vessels. To improve convergence during training, we apply a training approach that uses roll-back to reject stochastic changes to the network that increase the loss function. The trained neural network approximation is about 1000 times faster than the direct calculation for individual replicas. Because different applications will have different criteria for acceptable approximation accuracy, we discuss a variety of loss functions and accuracy estimators that can help select the best network for a particular application. We briefly discuss some of the issues we encountered with overfitting, mismapping of the field values and the geometrical conditions that lead to large absolute and relative errors in the approximate solution.

60 APPLIED LIFE SCIENCES↗

An Investigation of LES Wall Modeling for Rayleigh–Bénard Convection via Interpretable and Physics-Aware Feedforward Neural Networks with DNS

Abstract The traditional approach of using the Monin–Obukhov similarity theory (MOST) to model near-surface processes in large-eddy simulations (LESs) can lead to significant errors in natural convection. In this study, we propose an alternative approach based on feedforward neural networks (FNNs) trained on output from direct numerical simulation (DNS). To evaluate the performance, we conduct both a priori and a posteriori tests. In the a priori (offline) tests, we compare the statistics of the surface shear stress and heat flux, computed from filtered DNS input variables, to the stress and flux obtained from the filtered DNS. Additionally, we investigate the importance of various input features using the Shapley additive explanations value and the conditional average of the filter grid cells. In the a posteriori (online) tests, we implement the trained models in the System for Atmospheric Modeling (SAM) LES and compare the LES-generated surface shear stress and heat flux with those in the DNS. Our findings reveal that vertical velocity, a traditionally overlooked flow quantity, is one of the most important input features for determining the wall fluxes. Increasing the number of input features improves the a priori test results but does not always improve the model performance in the a posteriori tests because of the differences in input variables between the LES and DNS. Last, we show that physics-aware FNN models trained with logarithmic and scaled parameters can well extrapolate to more intense convection scenarios than in the training dataset, whereas those trained with primitive flow quantities cannot. Significance Statement The traditional near-surface turbulence model, based on a shear-dominated boundary layer flow, does not represent near-surface turbulence in natural convection. Using a feedforward neural network (FNN), we can construct a more accurate model that better represents the near-surface turbulence in various flows and reveals previously overlooked controlling factors and process interactions. Our study shows that the FNN-generated models outperform the traditional model and highlight the importance of the near-surface vertical velocity. Furthermore, the physics-aware FNN models exhibit the potential to extrapolate to convective flows of various intensities beyond the range of the training dataset, suggesting their broader applicability for more accurate modeling of near-surface turbulence.

54 ENVIRONMENTAL SCIENCES↗

Beyond Price Taker: Conceptual Design and Optimization of Integrated Energy Systems Using Machine Learning Market Surrogates

Future electricity generation systems must be optimized to provide flexibility that counteracts the variability of non-dispatchable renewable energy sources and ensures the reliability and safety of critical infrastructure, including the electric grid. The current state-of-the-art is to co-optimize the design and operation of integrated energy systems (IES) treating historical or predicted time-series electricity prices as fixed parameters. Recent literature has shown the limitations of this price taker assumption, which neglects how IES optimization decisions influence market outcomes. As such, this paper proposes a new optimization formulation that uses machine learning surrogate models, trained from a library of annual market operation simulations, to embed IES market interactions into the co-optimization problem directly. Using a thermal generator example built in the open-source IDAES computational environment, we show that the price taker approach routinely over-predicts annual revenues by 8% or more compared to a validation simulation, where the proposed approach has a typical relative error of 1% or less.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Calibration of the Diffusivity Predictions of Centipede Using Approximate Bayesian Computation and Applications in Nyx (Engineering Scale) and Xolotl-MARMOT (Meso-Scale) Simulations

Fission gas evolution and release in UO 2 nuclear fuel are important fuel performance metrics and occur in several distinct stages: 1) nucleation, growth and resolution of intra-granular bubbles, 2) diffusion to grain boundaries and 3) nucleation and growth of bubbles at grain boundaries, which eventually form a connected network (percolation) enabling release of gas from grain boundaries through connections to triple junctions, grain edges or free surfaces. The NE-SciDAC project is developing several computational tools to model this problem, which are connected in a hierarchical multi-scale framework. The information transfer in the multi-scale framework is a critical step that, in addition to best-estimates, should include uncertainty quantification. Despite taking a first-principles multi-scale approach, there is a need to perform parameter calibration to ensure consistency with available experimental data. In the present study, uncertainty quantification (UQ) and parameter calibration is demonstrated for one of the lower length scale codes in the multi-scale framework (Centipede) and then the results, including instances of the propagated uncertainties, are used in other codes within the framework, specifically Nyx and Xolotl-MARMOT. We calibrated the model parameters in Centipede, a computer code used to predict diffusivities of uranium (U) and xenon (Xe) in the context of the simulation of fission gas in uranium oxide (UO 2 ) nuclear fuel. The Centipede code depends on 183 parameters, all of which are subject to uncertainty. The three data sets used in our calibration effort are taken from the literature. This data is available as a set of measurements, including measurement errors. Our goal is to calibrate a statistical model that predicts both the value of the measurement and the uncertainty associated with the measurement. We perform a Bayesian calibration of the model parameters using a dedicated approximate Bayesian computation (ABC) likelihood function. To avoid excessive computational costs, we replace the expensive Centipede simulation code by a higher-order surrogate model, constructed using only the 9 most important parameters. These important parameters are identified by a preliminary global sensitivity analysis (GSA) study. Among the important parameters are T0 (the temperature at which UO 2 is perfectly stoichiometric) and Hf_pO2 (the temperature dependence of the oxygen (O) partial pressure) that should be considered as operating conditions to be estimated along with the other parameters. We consider two different cases: one where we define one set of these operating conditions for all data sets, and one where we define distinct operating condition parameters for each data set. The Xe diffusivities predicted by the latter case show distinct features that could not be observed in the former. Next, we use the diffusivity predictions by Centipede as input to Nyx, a reduced order fuel performance code focused on gas behavior alone, in order to estimate quantities associated with inter-granular bubble formation at conditions specified by the experiments. Finally, the diffusivities obtained from the calibrated Centipede runs were used in coupled Xolotl-MARMOT simulations of intra- and inter-granular gas evolution. The results are compared to simulations using the baseline diffusivities from Turnbull et al.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Optical measure of disorder: Why Urbach analysis works for amorphous silicon but fails for amorphous carbon

We investigate the basis for the Paradox where the Urbach-slope disorder parameter disagrees with the Raman width in amorphous diamond-like carbon (DLC) materials (lower Urbach-Eo yet greater Raman width). We examined the bandgap and Urbach-slope measurement issues. While significant errors are identified, ultimately these cannot resolve the Paradox. This resolution involves large changes in shape of DLC's absorption band resulting from large energy shifts. To solidify the understanding of band energy shifts, we examined their properties using a-Si:H, well known for its agreement between Raman and Urbach-slope disorder parameters. Using Tauc-Lorentz (T-L) dispersion functions, the Urbach method works for a-Si:H because small changes in shape occur, as measured by (i) peak position, (ii) peak width, (iii) peak amplitude, and (iv) and gap energies. Moreover, for a-Si:H, a 5th condition (v) there is no other feature below the principal band to interfere with its fall. All 5 of these conditions fail for DLCs. Because T-L dispersions accommodate asymmetry, we traced the Paradox to a steepening of the band shape as it pushes toward lower energy. This steepening is unrelated to disorder, but to a boundary value problem where energy cannot be negative. We conclude that Urbach analysis is invalid for DLC materials, but the T-L width is a good measure instead. We also examined another disorder measure, where π bonds should be considered as disorder relative to the σ bonds in an idealized “amorphous diamond” sp 3 network: superconvergence can be used to quantify these π transitions. Finally, further molecular orbital calculations are needed to properly interpret the changes in T-L widths.

36 MATERIALS SCIENCE↗

Deep learning based x-ray spectrometer for high repetition rate characterization of betatron radiation

Betatron radiation produced from a laser-wakefield accelerator is a broadband, hard x-ray (>1 keV) source that has been used in a variety of applications in medicine, engineering, and fundamental science. Further development and optimization of stable, high repetition rate (HRR) (>1 Hz) betatron sources will provide a means to extend their application base to include single-shot dynamical measurements of ultrafast processes or dense materials. Recent advances in laser technology used in such experiments have enabled increases in shot-rate and system stability, providing improved statistical analysis and detailed parameter scans. However, unique challenges exist at high repetition rate, where data throughput and source optimization are now limited by diagnostic acquisition rates and analysis. Here, we present the development of a machine-learning algorithm for the real-time analysis of betatron radiation. We report on the fielding of this deep learning algorithm for online source characterization at the Institut National de la Recherche Scientifique's Advanced Laser Light Source. By fine-tuning an algorithm originally trained on a fully synthetic dataset using a subset of experimental data, the algorithm can predict the betatron critical energy with a percent error of 7.2 % with a reconstruction time of 1.5 ms, providing a valuable tool for real-time, multi-objective optimization at HRR.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Towards Generalizable and Efficient Circuit Topology Design: A Graph-Transformer-based Surrogate Model with Curriculum Learning

Unlike circuit parameter and sizing optimizations, the automated design of analog circuit topologies poses significant challenges for learning-based approaches. One challenge arises from the combinatorial growth of the topology space with circuit size, which limits the topology optimization efficiency. Moreover, traditional circuit evaluation methods are time-consuming, while the presence of data discontinuity in the topology space makes the accurate prediction of circuit performance exceptionally difficult for unseen topologies. To tackle these challenges, we design a novel Graph-Transformer-based Network (GTN) as the surrogate model for circuit evaluation, offering a substantial acceleration in the speed of circuit topology optimization without sacrificing performance. Our GTN model architecture is designed to embed voltage changes in circuit loops and current flows in connected devices, enabling accurate performance predictions for circuits with unseen topologies. To address the cold start problem when scaling GTN to large-scale circuits, we further introduce a curriculum learning strategy that progressively trains GTN from small-scale to large-scale circuits. This approach enables the model to first learn fundamental physical principles from simpler topologies and gradually adapt to complex configurations, effectively bridging the circuit complexity gap and improving prediction accuracy. Taking the power converter circuit design as an experimental task, our GTN model significantly outperforms an analytical approach and baseline methods directly utilizing graph neural networks. Furthermore, GTN achieves less than 5% relative error and 196× speed-up compared with high-fidelity simulation. Notably, our GTN surrogate model empowers an automatic circuit design framework to discover circuits of comparable quality to those identified through high-fidelity simulation while reducing the time required by up to 98.2%. With curriculum learning, the enhanced GTN achieves a 51% improvement for performance prediction of large-scale circuits compared to the GTN model without this strategy. These advancements establish GTN as a scalable framework for automated analog circuit design across varying circuit complexity levels.

Lu, Haoshu [New Jersey Institute of Technology (NJ↗

Improving Subsurface Stress Characterization for Carbon Dioxide Storage Projects by Incorporating Machine Learning Techniques

The overall objective of this project is to develop a framework for reliable characterization and prediction of the state of stress in the overburden and underburden (including the basement) in CO 2 storage reservoirs using machine learning and integrated geomechanics and geophysical methods. Specifically, we propose to develop workflow encompassing of technologies and/or methods to predict stress and pressure changes due to CO 2 injection in an active tertiary recovery site and their impacts on subtle fault activation, fractures and occurrence of microseismic events and compare responses to field observations. In this project, we anticipate using dataset from the Farnsworth field Unit (FWU) which is operated by Purdure Petroleum. A novel elastic-waveform VSP inversion technique will be used to estimate high-resolution spatial and temporal changes of elastic moduli in CO 2 storage reservoirs, which will be combined with velocity-stress relationship derived from laboratory tests to obtain subsurface pressure and stress. Clustered microseismic data will be jointly inverted for improved focal mechanisms. Least-squares reverse-time migration of microseismic waveform data will be performed to directly image fracture/fault zones. Additionally, a deep neural network machine learning technique with convolutional and recurrent layers will be used for learning the spectro-temporal structures in microseismic waveforms. The results of this geotechnical data analysis will be integrated to develop a high-resolution 3D mechanical earth model extending from the overburden sealing formations to the underburden including the basement. Mechanical properties will be derived through integration of mechanical logs, tests, available results from chemo-mechanical laboratory tests, and elastic inversion of seismic data using a combination of Bayesian and stochastic methods as well as machine learning technique. Failure features (faults/fractures) will be represented and/or modeled based on seismic and core data analysis. A transient hydrodynamic-geomechanical model will be developed through coupling with the calibrated FWU reservoir simulation model. The full physics coupled model will be used to train a reduced order proxy model using machine learning algorithm for estimating stress which will then be used with appropriate constitutive relationships and forward seismological models to simulate pressure changes and induced microseismicity. An advanced optimization framework will be developed to perform a history match to minimize error between field observations and simulated. The history matched proxy model will be verified against the full-physics equivalent. The field observations that will be used in the coupled model calibration process include pressure/stress inverted from VSP, moment magnitude from microseismic analysis, real time downhole pressure measurements, production and injection data. Parameter sensitivity and uncertainty analysis will be performed to characterize the impact of model parameter uncertainty on stress estimates. The proposed project will have significant impact on future field implementation of the proposed technology. Because the project field site is an ongoing CO 2 EOR development, the value of the new technology will be demonstrated in an operational context and evaluated as a viable risk mitigation strategy. Cost/benefit will be evaluated together with the various commercial incentives for CO 2 sequestration available to oil and gas operators. The extensive available dataset and ongoing data acquisition under the SWP Phase III work plan provides flexibility for investigation of multiple approaches and reduces technical risk.

58 GEOSCIENCES↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

NSFnets (Navier-Stokes flow nets): Physics-informed neural networks for the incompressible Navier-Stokes equations

In the last 50 years there has been a tremendous progress in solving numerically the Navier-Stokes equations using finite differences, finite elements, spectral, and even meshless methods. Yet, in many real cases, we still cannot incorporate seamlessly (multi-fidelity) data into existing algorithms, and for industrial-complexity applications the mesh generation is time consuming and still an art. Moreover, solving ill-posed problems (e.g., lacking boundary conditions) or inverse problems is often prohibitively expensive and requires different formulations and new computer codes. Here, we employ physics-informed neural networks (PINNs), encoding the governing equations directly into the deep neural network via automatic differentiation, to overcome some of the aforementioned limitations for simulating incompressible laminar and turbulent flows. We develop the Navier-Stokes flow nets (NSFnets) by considering two different mathematical formulations of the Navier-Stokes equations: the velocity-pressure (VP) formulation and the vorticity-velocity (VV) formulation. Since this is a new approach, we first select some standard benchmark problems to assess the accuracy, convergence rate, computational cost and flexibility of NSFnets; analytical solutions and direct numerical simulation (DNS) databases provide proper initial and boundary conditions for the NSFnet simulations. The spatial and temporal coordinates are the inputs of the NSFnets, while the instantaneous velocity and pressure fields are the outputs for the VP-NSFnet, and the instantaneous velocity and vorticity fields are the outputs for the VV-NSFnet. This is unsupervised learning and, hence, no labeled data are required beyond boundary and initial conditions and the fluid properties. The residuals of the VP or VV governing equations, together with the initial and boundary conditions, are embedded into the loss function of the NSFnets. No data is provided for the pressure to the VP-NSFnet, which is a hidden state and is obtained via the incompressibility constraint without extra computational cost. Unlike the traditional numerical methods, NSFnets inherit the properties of neural networks (NNs), hence the total error is composed of the approximation, the optimization, and the generalization errors. Here, we empirically attempt to quantify these errors by varying the sampling (“residual”) points, the iterative solvers, and the size of the NN architecture. For the laminar flow solutions, we show that both the VP and the VV formulations are comparable in accuracy but their best performance corresponds to different NN architectures. The initial convergence rate is fast but the error eventually saturates to a plateau due to the dominance of the optimization error. For the turbulent channel flow, we show that NSFnets can sustain turbulence at , but due to expensive training we only consider part of the channel domain and enforce velocity boundary conditions on the subdomain boundaries provided by the DNS data base. We also perform a systematic study on the weights used in the loss function for balancing the data and physics components, and investigate a new way of computing the weights dynamically to accelerate training and enhance accuracy. In the last part, we demonstrate how NSFnets should be used in practice, namely for ill-posed problems with incomplete or noisy boundary conditions as well as for inverse problems. We obtain reasonably accurate solutions for such cases as well without the need to change the NSFnets and at the same computational cost as in the forward well-posed problems. As a result, we also present a simple example of transfer learning that will aid in accelerating the training of NSFnets for different parameter settings.

97 MATHEMATICS AND COMPUTING↗

Cosmological Constraints from Combining Photometric Galaxy Surveys and Gravitational Wave Observatories

Spatial variations in survey properties due to selection effects generate substantial systematic errors in large-scale structure measurements in optical galaxy surveys on very large scales. On such scales, the statistical sensitivity of optical surveys is also limited by their finite sky coverage. By contrast, gravitational wave (GW) sources appear to be relatively free of these issues, provided the angular sensitivity of GW experiments can be accurately characterized. We quantify the expected cosmological information gain from combining the forecast LSST 3$\times$2pt analysis (combination of three 2-point correlations of galaxy density and weak lensing shear fields) with the large-scale auto-correlation of GW sources from proposed next-generation GW experiments. We find that in $\Lambda$CDM and $w$CDM models, there is no significant improvement in cosmological constraints from combining GW with LSST 3$\times$2pt over LSST alone, due to the large shot noise for the former; however, this combination does enable a $\sim6\%$ constraint on the linear galaxy bias of GW sources. More interestingly, the optical-GW data combination provides tight constraints on models with primordial non-Gaussianity (PNG), due to the predicted scale-dependent bias in PNG models on large scales. Assuming that the largest angular scales that LSST will probe are comparable to those in Stage III surveys ($\ell_{\rm min}\sim50$), the inclusion of next-generation GW measurements could improve constraints on the PNG parameter $f_{\rm NL}$ by up to a factor of $\simeq6.6$ compared to LSST alone, yielding $\sigma(f_{\rm NL})=8.5$. These results assume the expected capability of a network of Einstein Telescope-like GW observatories, with a detection rate of $10^6$ events/year. We investigate the sensitivity of our results to different assumptions about future GW detectors as well as different LSST analysis choices.

79 ASTRONOMY AND ASTROPHYSICS↗