Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

LTAU-FF: Loss Trajectory Analysis for Uncertainty in atomistic Force Fields

Model ensembles are effective tools for estimating prediction uncertainty in deep learning atomistic force fields. However, their widespread adoption is hindered by high computational costs and overconfident error estimates. In this work, we address these challenges by leveraging distributions of per-sample errors obtained during training and employing a distance-based similarity search in the model latent space. Our method, which we call LTAU (Loss Trajectory Analysis for Uncertainty), efficiently estimates the full probability distribution function of errors for any test point using the logged training errors, achieving speeds that are 2–3 orders of magnitudes faster than typical ensemble methods and allowing it to be used for tasks where training or evaluating multiple models would be infeasible. We apply LTAU towards estimating parametric uncertainty in atomistic force fields (LTAU-FF), demonstrating that it produces well-calibrated confidence intervals and predicts errors that correlate strongly with the true errors for data near the training domain. Furthermore, we show that the errors predicted by LTAU-FF can be used in practical applications for detecting out-of-domain data, tuning model performance, and predicting failure during simulations. We believe that LTAU will be a valuable tool for uncertainty quantification in atomistic force fields and is a promising method that should be further explored in other domains of machine learning.

97 MATHEMATICS AND COMPUTING↗

Realistic simulations of spin squeezing and cooperative coupling effects in large ensembles of interacting two-level systems

We describe an efficient numerical method for simulating the dynamics of interacting spin ensembles in the presence of dephasing and decay. The method builds on the discrete truncated Wigner approximation for isolated systems, which combines the mean-field dynamics of a spin ensemble with a Monte Carlo sampling of discrete initial spin values to account for quantum correlations. Here we show how this approach can be generalized for dissipative spin systems by replacing the deterministic mean-field evolution by a stochastic process, which describes the decay of coherences and populations while preserving the length of each spin. We demonstrate the application of this technique for simulating nonclassical spin-squeezing effects or the dynamics and steady states of cavity QED models with 10 5 interacting two-level systems. This opens up the possibility to perform accurate real-scale simulations of a diverse range of experiments in quantum optics or with solid-state spin ensembles under realistic laboratory conditions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Updating and Evaluating Anthropogenic Emissions for NOAA’s Global Ensemble Forecast Systems for Aerosols (GEFS-Aerosols): Application of an SO 2 Bias-Scaling Method

We updated the anthropogenic emissions inventory in NOAA’s operational Global Ensemble Forecast for Aerosols (GEFS-Aerosols) to improve the model’s prediction of aerosol optical depth (AOD). We used a methodology to quickly update the pivotal global anthropogenic sulfur dioxide (SO 2 ) emissions using a speciated AOD bias-scaling method. The AOD bias-scaling method is based on the latest model predictions compared to NASA’s Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA2). The model bias was subsequently applied to the CEDS 2019 SO 2 emissions for adjustment. The monthly mean GEFS-Aerosols AOD predictions were evaluated against a suite of satellite observations (e.g., MISR, VIIRS, and MODIS), ground-based AERONET observations, and the International Cooperative for Aerosol Prediction (ICAP) ensemble results. The results show that transitioning from CEDS 2014 to CEDS 2019 emissions data led to a significant improvement in the operational GEFS-Aerosols model performance, and applying the bias-scaled SO 2 emissions could further improve global AOD distributions. The biases of the simulated AODs against the observed AODs varied with observation type and seasons by a factor of 3~13 and 2~10, respectively. The global AOD distributions showed that the differences in the simulations against ICAP, MISR, VIIRS, and MODIS were the largest in March–May (MAM) and the smallest in December–February (DJF). When evaluating against the ground-truth AERONET data, the bias-scaling methods improved the global seasonal correlation (r), Index of Agreement (IOA), and mean biases, except for the MAM season, when the negative regional biases were exacerbated compared to the positive regional biases. The effect of bias-scaling had the most beneficial impact on model performance in the regions dominated by anthropogenic emissions, such as East Asia. However, it showed less improvement in other areas impacted by the greater relative transport of natural emissions sources, such as India. The accuracies of the reference observation or assimilation data for the adjusted inputs and the model physics for outputs, and the selection of regions with less seasonal emissions of natural aerosols determine the success of the bias-scaling methods. A companion study on emission scaling of anthropogenic absorbing aerosols needs further improved aerosol prediction.

54 ENVIRONMENTAL SCIENCES↗

Be It Unresolved: Measuring Time Delays from Lensed Supernovae

Gravitationally lensed Type Ia supernovae (SNe Ia) may be the next frontier in cosmic probes, able to deliver independent constraints on dark energy, spatial curvature, and the Hubble constant. Measurements of time delays between the multiple images become more incisive due to the standardized candle nature of the source, monitoring for months rather than years, and partial immunity to microlensing. While currently extremely rare, hundreds of such systems should be detected by upcoming time domain surveys. Others will have the images spatially unresolved, with the observed lightcurve a superposition of time-delayed image fluxes. We investigate whether unresolved images can be recognized as lensed sources given only lightcurve information, and whether time delays can be extracted robustly. We develop a method that we show can identify these systems for the case of lensed SNe Ia with two images and time delays exceeding ten days. When tested on such an ensemble, without microlensing, the method achieves a false-positive rate of ≲5%, and measures the time delays with a completeness of ≳93% and with a bias of ≲0.5% for Δt fit ≳ 10 days. Since the method does not assume a template of any particular type of SN, the method has the (untested) potential to work on other types of lensed SNe systems and possibly on other transients as well.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

Computational Response Theory for Dynamics

Quantifying the sensitivity - how a quantity of interest (QoI) varies with respect to a parameter – and response – the representation of a QoI as a function of a parameter - of a computer model of a parametric dynamical system is an important and challenging problem. Traditional methods fail in this context since sensitive dependence on initial conditions implies that the sensitivity and response of a QoI may be ill-conditioned or not well-defined. If a chaotic model has an ergodic attractor, then ergodic averages of QoIs are well-defined quantities and their sensitivity can be used to characterize model sensitivity. The response theorem gives sufficient conditions such that the local forward sensitivity – the derivative with respect to a given parameter - of an ergodic average of a QoI is well-defined. We describe a method based on ergodic and response theory for computing the sensitivity and response of a given QoI with respect to a given parameter in a chaotic model with an ergodic and hyperbolic attractor. This method does not require computation of ensembles of the model with perturbed parameter values. The method is demonstrated and some of the computations are validated on the Lorenz 63 and Lorenz 96 models.

97 MATHEMATICS AND COMPUTING↗

Generative adversarial networks for ensemble projections of future urban morphology

As city planners design and adapt cities for future resilience and intelligence, interactions among neighborhood morphological development with respect to changes in population and resultant built infrastructure's impact on the natural environment must be considered. For deep understanding of these interactions, explicit representation of future neighborhoods is necessary for future city modeling. Generative Adversarial Networks (GANs) have been shown to produce spatially accurate urban forms at scales representing entire cities to those at neighborhood and single building scale. Here we demonstrate a GAN method for generating an ensemble of possible new neighborhoods given land use characteristics and designated neighborhood type.

Dumas, Melissa↗

Enhancing Automotive Intrusion Detection Through Multi-Modal Fusion: A CAN FD-LiDAR Approach

As vehicles become smarter and more autonomous, they increasingly depend on advanced sensors and communication technologies to operate securely. However, such growing dependence on technology—whether it’s CAN (Controller Area Network) for internal communication or LiDAR (Light Detection and Ranging) for sensing the world around them—also expands the attack surface for the types of cyber attacks. Traditional intrusion detection systems (IDS) typically monitor these systems in isolation, limiting their ability to detect sophisticated, crosssystem attacks. To address this, we propose a multi-modal fusion approach that combines real-world CAN FD signals (from the HCRL dataset) with LiDAR features (from the nuScenes dataset) to enhance attack detection. Our method employs a twostage ensemble approach. Calibrated XGBoost and LightGBM models initially process CAN FD (Fuzzing Data) and LiDAR data independently, detecting timing anomalies and space abnormalities. They are subsequently logarithmically combined with a logistic regression meta-model along with 17 engineered features capturing cross-modal behavior, prediction conflicts, and nonlinear interactions. This approach achieves an AUC of 0.87 and an F1-score of 0.82, surpassing single-modality baselines and early fusion methods, at merely 2 ms inference latency. Compared with deep learning competitors, it is 3 times more efficient, providing a lightweight, interpretable, and real time solution to automotive cybersecurity.

97 MATHEMATICS AND COMPUTING↗

Generative Ensemble Regression: Learning Particle Dynamics from Observations of Ensembles with Physics-informed Deep Generative Models

Here, we propose a new method for inferring the governing stochastic ordinary differential equations (SODEs) by observing particle ensembles at discrete and sparse time instants, i.e., multiple “snapshots.” Particle coordinates at a single time instant, possibly noisy or truncated, are recorded in each snapshot but are unpaired across the snapshots. By training a physics-informed generative model that generates “fake” sample paths, we aim to fit the observed particle ensemble distributions with a curve in the probability measure space, which is induced from the inferred particle dynamics. We employ different metrics to quantify the differences between distributions, e.g., the sliced Wasserstein distances and the adversarial losses in generative adversarial networks. We refer to this method as generative “ensemble-regression” (GER), in analogy to the classic “point-regression,” where we infer the dynamics by performing regression in the Euclidean space. We illustrate the GER by learning the drift and diffusion terms of particle ensembles governed by SODEs with Brownian motions and Lévy processes up to 100 dimensions. We also discuss how to treat cases with noisy or truncated observations. Apart from systems consisting of independent particles, we also tackle nonlocal interacting particle systems with unknown interaction potential parameters by constructing a physics-informed loss function. Finally, we investigate scenarios of paired observations and discuss how to reduce the dimensionality in such cases by proving a convergence theorem that provides theoretical support.

97 MATHEMATICS AND COMPUTING↗

Designing robust energy policy packages under deep uncertainty: A multi-metric decision support framework

The complexity of transitioning to sustainable energy systems requires policy frameworks capable of balancing multiple objectives while addressing deep uncertainty. However, existing approaches often lack systematic methods to identify combinations of policy levers that remain effective across a wide range of uncertain futures. This paper presents a novel decision support framework that guides the selection of robust policy packages based on their performance across multiple objectives under uncertainty. Our method leverages a large ensemble of scenarios and applies scenario discovery techniques to identify influential policy levers. Here, we introduce new indicators to assess the robustness of policies by evaluating their ability to mitigate adverse outcomes across metrics. These indicators support an iterative process to build a robust policy package. Finally, we map the technological and energy pathways associated with the robust policy package by leveraging an energy system optimization model. We illustrate the application of this framework to the Spanish energy system, providing insights into how specific combinations of policy levers shape decarbonization pathways under uncertainty.

Decision-support method↗

Performance assessment of ensembles of in situ workflows under resource constraints

Summary Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long‐running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble‐based simulations paired with in situ analyses using in‐memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co‐location scenarios, our proposed performance indicators demonstrate benefits of co‐locating simulation and coupled analyses within a compute node.

Do, Tu Mai Anh↗

A deep learning-accelerated data assimilation and forecasting workflow for commercial-scale geologic carbon storage

Fast assimilation of monitoring data to update forecasts of pressure buildup and carbon dioxide (CO 2 ) plume migration under geologic uncertainties is a challenging problem in geologic carbon storage. The high computational cost of data assimilation with a high-dimensional parameter space impedes fast decision-making for commercial-scale reservoir management. We propose to leverage physical understandings of porous medium flow behavior with deep learning techniques to develop a fast data assimilation-reservoir response forecasting workflow. Applying an Ensemble Smoother Multiple Data Assimilation (ES-MDA) framework, the workflow updates geologic properties and predicts reservoir performance with quantified uncertainty from pressure history and CO 2 plumes interpreted through seismic inversion. As the most computationally expensive component in such a workflow is reservoir simulation, we developed surrogate models to predict dynamic pressure and CO 2 plume extents under multi-well injection. The surrogate models employ deep convolutional neural networks, specifically, a wide residual network and a residual U-Net. The workflow is validated against a flat threedimensional reservoir model representative of a clastic shelf depositional environment. Intelligent treatments are applied to bridge between quantities in a true-3D reservoir model and those in a single-layer reservoir model. The workflow can complete history matching and reservoir forecasting with uncertainty quantification in less than one hour on a mainstream personal workstation.

25 ENERGY STORAGE↗

Improving Prediction of Surface Solar Irradiance Variability by Integrating Observed Cloud Characteristics and Machine Learning

A 5-year, 1-minute resolution observational dataset of clouds and solar radiation was produced that includes two metrics of the variability in surface solar irradiance due to cloud type and fractional sky cover. Multiple regression models were trained to fit observations of surface solar irradiance variability from those two cloud property predictors. We found that ensemble tree-based methods, Random Forest and Gradient Boosting Machine, have the least overfitting issues and showed the best performance with an R2 of 0.42. While the observational data trained in this study was only from one site, the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site in Oklahoma, initial comparisons of the seasonality of the statistics suggest that these results are relatively weather regime independent; the generality of such a finding across sites will be tested in future work. The observational data and developed machine learning model are being used to create a numerical weather prediction model parameterization to enable day-ahead solar variability prediction in a computationally efficient way. This is a first step towards creating a new paradigm of predicting day-ahead variability with the potential to provide a new tool to improve grid operation, planning, and resilience.

Riihimaki, Laura↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Fluctuations in heavy ion collisions and global conservation effects

Subensemble is a type of statistical ensemble which is the generalization of grand canonical and canonical ensembles. The subensemble acceptance method (SAM) provides general formulas to correct the cumulants of distributions in heavy-ion collisions for the global conservation of all QCD charges. The method is applicable for an arbitrary equation of state and sufficiently large systems, such as those created in central collisions of heavy ions. The new fluctuation measures insensitive to global conservation effects are presented. The main results are illustrated in the hadron resonance gas and van der Waals fluid frameworks.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Antarctic ice sheet model comparison with uncurated geological constraints shows that higher spatial resolution improves deglacial reconstructions

Accurately reconstructing past changes to the shape and volume of the Antarctic ice sheet relies on the use of physically based and thus internally consistent ice sheet modeling, benchmarked against spatially limited geologic data. The challenge in model benchmarking against geologic data is diagnosing whether model-data misfits are the result of an inadequate model, inherently noisy or biased geologic data, and/or incorrect association between modeled quantities and geologic observations. In this work we address this challenge by (i) the development and use of a new model-data evaluation framework applied to an uncurated data set of geologic constraints, and (ii) nested high-spatial-resolution modeling designed to test the hypothesis that model resolution is an important limitation in matching geologic data. While previous approaches to model benchmarking employed highly curated datasets, our approach applies an automated screening and quality control algorithm to an uncurated public dataset of geochronological observations (specifically, cosmogenic-nuclide exposure-age measurements from glacial deposits in ice-free areas). This optimizes data utilization by including more geological constraints, reduces potential interpretive bias, and allows unsupervised assimilation of new data as they are collected. We also incorporate a nested model framework in which high-resolution domains are downscaled from a continent-wide ice sheet model. We highlight the application of this framework by applying these methods to a small ensemble of deglacial ice-sheet model simulations, and demonstrate that the nested approach improves the ability of model simulations to match exposure age data collected from areas of complex topography and ice flow. We develop a range of diagnostic model-data comparison metrics to provide more insight into model performance than possible from a single-valued misfit statistic, showing that different metrics capture different aspects of ice sheet deflation.

Geosciences↗

An Ensemble Score Filter for Tracking High-Dimensional Nonlinear Dynamical Systems

We propose an ensemble score filter (EnSF) for solving high-dimensional nonlinear filtering problems with superior accuracy. A major drawback of existing filtering methods, e.g., particle filters or ensemble Kalman filters, is the low accuracy in handling high-dimensional and highly nonlinear problems. EnSF attacks this challenge by exploiting the score-based diffusion model, defined in a pseudo-temporal domain, to characterizing the evolution of the filtering density. EnSF stores the information of the recursively updated filtering density function in the score function, in stead of storing the information in a set of finite Monte Carlo samples (used in particle filters and ensemble Kalman filters). Unlike existing diffusion models that train neural networks to approximate the score function, we develop a training-free score estimation that uses mini-batch-based Monte Carlo estimator to directly approximate the score function at any pseudo-spatial-temporal location, which provides sufficient accuracy in solving high-dimensional nonlinear problems as well as saves tremendous amount of time spent on training neural networks. Another essential aspect of EnSF is its analytical update step, gradually incorporating data information into the score function, which is crucial in mitigating the degeneracy issue faced when dealing with very high-dimensional nonlinear filtering problems. High-dimensional Lorenz systems are used to demonstrate the performance of our method. EnSF provides surprisingly impressive performance in reliably tracking extremely high-dimensional Lorenz systems (up to 1,000,000 dimension) with highly nonlinear observation processes, which is a well-known challenging problem for existing filtering methods.

97 MATHEMATICS AND COMPUTING↗