Engineering PapersSearch

SEARCH · Engineering Papers

Results for “ensembles”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur

Huge ensembles – Part 2: Properties of a huge ensemble of hindcasts generated with spherical Fourier neural operators

Abstract. In Part 1, we created an ensemble based on spherical Fourier neural operators. As initial condition perturbations, we used bred vectors, and as model perturbations, we used multiple checkpoints trained independently from scratch. Based on diagnostics that assess the ensemble's physical fidelity, our ensemble has comparable performance to operational weather forecasting systems. However, it requires orders-of-magnitude fewer computational resources. Here in Part 2, we generate a huge ensemble (HENS), with 7424 members initialized each day of summer 2023. We enumerate the technical requirements for running huge ensembles at this scale. HENS precisely samples the tails of the forecast distribution and presents a detailed sampling of internal variability. HENS has two primary applications: (1) as a large dataset with which to study the statistics and drivers of extreme weather and (2) as a weather forecasting system. For extreme climate statistics, HENS samples events 4σ away from the ensemble mean. At each grid cell, HENS increases the skill of the most accurate ensemble member and enhances coverage of possible future trajectories. As a weather forecasting model, HENS issues extreme weather forecasts with better uncertainty quantification. It also reduces the probability of outlier events, in which the verification value lies outside the ensemble forecast distribution.

Mahesh, Ankur

From Ensemble Climate to Ensemble Impacts

Many climate-risk tools rely on ensemble mean projections or endpoint climate snapshots to characterize future hazards. Although convenient for communication, these representations remove the statistical, temporal, and physical information that real infrastructure systems respond to. Infrastructure degradation and failure arise from extremes, sequences, cumulative stress, compound hazards, and nonlinear fragility relationships, none of which survive ensemble averaging or temporal compression. Power-system failure statistics and cascading failure models further show that infrastructure risk is dominated by tail events and path-dependent dynamics rather than by mean conditions. This paper demonstrates why ensemble mean or endpoint-only climate representations are mathematically and physically inconsistent with engineering-grade risk analysis. We outline a model-resolved, time-series-based workflow that preserves extremes, variability, and sequencing by propagating each climate-model realization independently through hazard formation, exposure, fragility, and cascading failure mechanisms. Taking the ensemble of impacts—rather than the ensemble of climate—provides a defensible, physically coherent foundation for infrastructure resilience planning, regulatory compliance, and long-term investment decisions.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE

An Alternative Ensemble Streamflow Prediction Approach Using Improved Subseasonal Precipitation Forecasts from the North America Multi-Model Ensemble Phase II

In this article, streamflow forecasting at a subseasonal time scale (10–30 days into the future) is important for various human activities. The ensemble streamflow prediction (ESP) is a widely applied technique for subseasonal streamflow forecasting. However, ESP’s reliance on the randomly resampled historical precipitation limits its predictive capability. Available dynamical subseasonal precipitation forecasts provide an alternative to the randomly resampled precipitation in ESP. Prior studies found the predictive performance of raw subseasonal precipitation forecast is limited in many regions such as the central south of the United States, which raises questions about its effectiveness in assisting streamflow forecasting. To further assess the hydrologic applicability of dynamical subseasonal precipitation forecasts, we test the subseasonal precipitation forecast from North America Multi-Model Ensemble Phase II (NMME-2) at four watersheds in the central south region of the United States. The subseasonal precipitation forecasts are postprocessed with bias correction and spatial disaggregation (BCSD) to correct bias and improve spatial resolution before replacing the randomly resampled precipitation in ESP for streamflow predictions. The performance of the resulting streamflow predictions is benchmarked with ESP. Evaluation is conducted using Kling–Gupta Efficiency (KGE), continuous ranked probability score (CRPS), probability of detection (POD), false alarm ratios (FARs), as well as reliability diagrams. Our results suggest that BCSD-corrected subseasonal precipitation forecasts lead to overall improved streamflow predictions due to added skills in winter and spring. Our results also suggest that BCSD-corrected subseasonal precipitation forecasts lead to improved predictions on the occurrence of high-percentile streamflow values above 75%. Overall, BCSD-corrected subseasonal precipitation has shown promising performance, highlighting its potential broader applications for river and flood forecasting.

54 ENVIRONMENTAL SCIENCES

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE

Atomic Ordering-Induced Ensemble Variation in Alloys Governs Electrocatalyst On/Off States

The catalytic behavior of a material is influenced by ensembles—the geometric configuration of atoms. Traditional approaches, mainly utilizing solid-solution alloys in electrocatalysis, have often overlooked the challenges posed by concurrent changes in the electronic structure (i.e. d-band center) when the composition is altered. Here, this study introduces a methodology that distinctly separates the geometric effects (i.e. ensembles) from the electronic structure. We compare the reactivity of compositionally identical, but structurally different Pd 3 Bi ordered intermetallic and solid-solution alloys. Remarkably, we find that Pd 3 Bi intermetallics display nearly no reactivity for the methanol oxidation (MOR), while their solid-solution counterparts have significant reactivity. This highlights a unique case where materials with identical chemical compositions demonstrate drastically different catalytic behavior underscoring the critical importance of ensembles in electrocatalysis. Specifically, Pd 3 Bi intermetallics form smaller ensembles (average coordination number: 4.5 ± 1.6) with almost no measurable MOR activity at room temperature, in contrast to the solid-solution Pd 3 Bi that exhibit larger ensembles (average coordination number: 6.8 ± 0.9) and considerable MOR reactivity (0.5 mA cm −2 Pd ). An ordered Pd 3 Bi alloy, with an intermediate ensemble size (average coordination number: 5.3 ± 1.2), displays moderate MOR activity (0.1 mA cm −2 Pd ), further confirming the direct correlation between ensemble size and catalytic activity. Notably, all Pd 3 Bi alloys maintain similar electronic structures, because the chemical composition of the alloys is fixed, indicating that the differences in reactivity are predominantly from changes to the ensemble size. Our findings offer an approach for precisely controlling catalytic activity through manipulating the geometric configuration of the atoms within an alloy, paving the way for more efficient catalyst design.

alloys

Discovering the Multisectoral Impacts of Global Energy Sector Outcomes Through Multiple Ensemble Aggregation Measures

Understanding complex human-Earth system interactions often involves analyzing large scenario ensembles that encompass a wide range of plausible futures. These ensembles often require aggregation to summarize information based on specific criteria or conditions. However, previous research using global change scenario ensembles has largely overlooked how the choice of aggregation method influences the interpretation of results. To address this gap, we leverage a large ensemble data set designed to capture broad energy system dynamics generated using the Global Change Analysis Model. We first explore how energy-related uncertainties are propagated to both global and regional water-energy-food sectors. We then conduct a rank correlation analysis across seven ensemble aggregation measures and demonstrate the need to consider multiple measures in global change scenarios. Our results suggest that global water and food sector outcomes in the 21st century vary widely depending on different scenario assumptions. The global energy productivity is projected to improve by the end of the century across all scenarios. Moreover, regions facing water scarcity challenges in 2100 do not always overlap with those facing extreme energy and food sector outcomes. Although rank correlations across seven aggregation measures are relatively stable across sectors, we identify cases where relying on a single measure leads to losing critical information in the full ensemble. Reliance on a single aggregation measure can distort the interpretation of global change scenario outcomes. Instead, adopting multiple ensemble aggregation measures provides a more holistic understanding of global change scenario ensembles.

Kim, Gijoo

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES

Evaluating Ensemble Predictions of South Asian Monsoon Low Pressure System Genesis

Abstract Synoptic-scale vortices known as monsoon low pressure systems (LPSs) frequently produce intense precipitation and hydrological disasters in South Asia, so accurately forecasting LPS genesis is crucial for improving disaster preparedness and response. However, the accuracy of LPS genesis forecasts by numerical weather prediction models has remained unknown. Here, we evaluate the performance of two global ensemble models—the U.S. Global Ensemble Forecast System (GEFS) and the Ensemble Prediction System of the European Centre for Medium-Range Weather Forecasts (ECMWF)—in predicting LPS genesis during the years 2021–22. The GEFS successfully predicted about half the observed LPS genesis events 1–2 days in advance; the ECMWF model captured an additional 10% of observed genesis events. Both models had a false alarm ratio (FAR) of around 50% for 1–2-day lead times. In both ensembles, the control run typically exhibited a higher probability of detection (POD) of observed events and a lower FAR compared to the perturbed ensemble members. However, a consensus forecast, in which genesis is predicted when at least 20% of ensemble members forecast LPS formation, had POD values surpassing those of the control run for all lead times. Moreover, probabilistic predictions of genesis over the Bay of Bengal, where most LPSs form, were skillful, with the fraction of ensemble members predicting LPS formation over a 5-day lead time approximating the observed frequency of genesis, without any adjustment or bias correction.

Suhas, D. L.

Ensemble Effects on Hydroxide Bond Dissociation Free Energies in Polyoxovanadate Clusters

Understanding structure-property relationships is foundational to numerous modern chemistries, such as proton-coupled electron transfer (PCET). However, an experimentally measured property is the result of the behavior from an ensemble of molecules. Neglecting ensemble effects, especially under complex chemical environments, may obfuscate these relationships and lead to discrepancies between theory and experiment. In this work, we demonstrate the impact of configurational entropy and local chemical environments on hydroxide bond dissociation free energies [BDFE- (O−H)] for a set of polyoxovanadate nanoclusters, at ambient conditions. The O−H bond strengths are investigated via density functional theory (DFT) coupled with statistical thermodynamic analysis and bilinear modeling, and compared with previous experimental results on the same systems, namely electrochemical solutions of: [V 6 O 13−x (OH) x (TRIOL R ) 2 ] −2 (x = 2, 4, 6; R = NO 2 , Me) and [V 6 O 11−x (OMe) 2 (OH) x (TRIOL NO 2 ) 2 ] −2 (x = 2, 4). Interestingly, we find that ensemble effects, even at room temperature, can account for a significant portion of the BDFE(O−H) trend with the degree of reduction via H atom binding, which cannot be fully captured by single-structure, static DFT calculations. Moreover, we find that the ensemble effects may be replicated statistically, requiring only enumeration of energetically accessible H-binding sites. With the ensemble effects resolved, we present a simple bilinear model to reconcile remaining biases between experiment and ensemble-informed theory, which corelate with clusterspecific electronic environment differences. The bilinear model achieves outstanding accuracy vs experiments with a root-mean squared error of 0.4 kcal/mol. Finally, based on the physicochemical characteristics of hydrogen interaction with polyoxometalates, we present a simple methodology that captures the BDFE(O−H) trend while dramatically reducing required DFT calculations by 98% and achieving accuracy within 1 kcal/mol. Overall, this work elucidates the roles and structural origins of configurational entropy and chemical effects on polyoxometalate hydroxide bond energies, with potential applicability to various atomically precise metal oxide systems. Importantly, it introduces models for rapid and highly accurate property calculations in connection with experiments.

Cluster chemistry

Quantification of regional net CO 2 flux errors in the Orbiting Carbon Observatory-2 (OCO-2) v10 model intercomparison project (MIP) ensemble using airborne measurements

Inverse model intercomparison projects (MIPs) provide a chance to assess the uncertainties in inversion estimates arising from various sources. However, accurately quantifying ensemble CO 2 flux errors remains challenging and often relies on the ensemble spread. This study proposes a method for quantifying the errors in regional net surface–atmosphere CO 2 flux estimates from models taken from the Orbiting Carbon Observatory-2 (OCO-2) v10 MIP by using independent airborne CO 2 measurements for the period 2015–2017. We first calculate the root mean square error (RMSE) between the ensemble mean of posterior CO 2 concentrations and airborne observations and then isolate the CO 2 concentration errors caused solely by the ensemble mean of posterior net fluxes by subtracting the observation, representation, and transport errors from seven regions. Our analysis reveals that the flux errors projected onto CO 2 space account for 55 %–85 % of the regional average RMSE over the 3 years, ranging from 0.88 to 1.91 ppm. In five regions, the error estimates based on observations exceed those computed from the ensemble spread of posterior fluxes by a factor of 1.33–1.93, implying an underestimation of the actual flux errors, while their magnitudes are comparable in two regions. The adjoint sensitivity analysis identifies that the underestimation of flux errors is prominent where the magnitudes of fossil fuel emissions exceed those of terrestrial-biosphere fluxes by a factor of 3–31 over the 3 years. This suggests the presence of systematic biases in the inversion estimates associated with errors in the prescribed fossil fuel emissions common to all models. Our study emphasizes the value of airborne measurements for quantifying regional errors in ensemble net CO 2 flux estimates.

54 ENVIRONMENTAL SCIENCES

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing

Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure

First-principles fusion plasma simulations are both compute and memory intensive, and CGYRO is no exception. The use of many HPC nodes to fit the problem in the available memory thus results in significant communication overhead, which is hard to avoid for any single simulation. That said, most fusion studies are composed of ensembles of simulations, so we developed a new tool, named XGYRO, that executes a whole ensemble of CGYRO simulations as a single HPC job. By treating the ensemble as a unit, XGYRO can alter the global buffer distribution logic and apply optimizations that are not feasible on any single simulation, but only on the ensemble as a whole. The main saving comes from the sharing of the collisional constant tensor structure, since its values are typically identical between parameter-sweep simulations. This data structure dominates the memory consumption of CGYRO simulations, so distributing it among the whole ensemble results in drastic memory savings for each simulation, which in turn results in overall lower communication overhead.

CGYRO

Large Ensemble Exploration of Global Energy Transitions Under National Emissions Pledges

Global climate goals require a transition to a deeply decarbonized energy system. Meeting the objectives of the Paris Agreement through countries' nationally determined contributions and long-term strategies represents a complex problem with consequences across multiple systems shrouded by deep uncertainty. Robust, large-ensemble methods and analyses mapping a wide range of possible future states of the world are needed to help policymakers design effective strategies to meet emissions reduction goals. This study contributes a scenario discovery analysis applied to a large ensemble of 5,760 model realizations generated using the Global Change Analysis Model. Eleven energy-related uncertainties are systematically varied, representing national mitigation pledges, institutional factors, and techno-economic parameters, among others. The resulting ensemble maps how uncertainties impact common energy system metrics used to characterize national and global pathways toward deep decarbonization. Results show globally consistent but regionally variable energy transitions as measured by multiple metrics, including electricity costs and stranded assets. Larger economies and developing regions experience more severe economic outcomes across a broad sampling of uncertainty. The scale of CO 2 removal globally determines how much the energy system can continue to emit, but the relative role of different CO 2 removal options in meeting decarbonization goals varies across regions. Previous studies characterizing uncertainty have typically focused on a few scenarios, and other large-ensemble work has not (to our knowledge) combined this framework with national emissions pledges or institutional factors. Our results underscore the value of large-ensemble scenario discovery for decision support as countries begin to design strategies to meet their goals.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Nanoscale wetting controls reactive Pd ensembles in synthesis of dilute PdAu alloy catalysts

The performance of bimetallic dilute alloy catalysts is largely determined by the size of minority metal ensembles on the nanoparticle surface. By analyzing the synthesis of catalysts comprising Pd 8 Au 92 nanoparticles supported on silica using surface-sensitive techniques, we report that whether Pd overgrowth occurs before or after Au nanoparticle deposition onto the support controls the surface Pd ensemble size and abundance. These differences in Pd ensembles influence catalytic reactivity in H 2 –D 2 isotope exchange and benzaldehyde hydrogenation, which, in correlation with theoretical calculations, is used to elucidate the active site(s) in each reaction. To clarify how the synthetic sequence controls the formation of Pd ensembles, we combine numerical wetting calculations and molecular dynamics simulations (with a machine-learned force field) to visualize Pd deposition and migration on the nanoparticle surface, respectively. Our results suggest that the nanoparticle–support interface restricts nanoparticle accessibility to Pd deposition, which consequently controls the Pd ensemble size, illustrating the critical role of nanoscale wetting phenomena during bimetallic catalyst preparation.

36 MATERIALS SCIENCE