Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Wave function methods for canonical ensemble thermal averages in correlated many-fermion systems

We present a wave function representation for the canonical ensemble thermal density matrix by projecting the thermofield double state against the desired number of particles. Furthermore, the resulting canonical thermal state obeys an imaginary-time evolution equation. Starting with the mean-field approximation, where the canonical thermal state becomes an antisymmetrized geminal power (AGP) wave function, we explore two different schemes to add correlation: by number-projecting a correlated grand-canonical thermal state and by adding correlation to the number-projected mean-field state. As benchmark examples, we use number-projected configuration interaction and an AGP-based perturbation theory to study the hydrogen molecule in a minimal basis and the six-site Hubbard model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty quantification for neural network potential foundation models

Abstract For neural network potentials (NNPs) to gain widespread use, researchers must be able to trust model outputs. However, the blackbox nature of neural networks and their inherent stochasticity are often deterrents, especially for foundation models trained over broad swaths of chemical space. Uncertainty information provided at the time of prediction can help reduce aversion to NNPs. In this work, we detail two uncertainty quantification (UQ) methods. Readout ensembling, by finetuning the readout layers of an ensemble of foundation models, provides information about model uncertainty, while quantile regression, by replacing point predictions with distributional predictions, provides information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to the foundation model and a series of finetuned models. The uncertainties produced by the readout ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.

36 MATERIALS SCIENCE↗

A method for examining ensemble averaging forms during the transition to turbulence in HED systems for application to RANS models

This paper discusses a strategy to initialize a two-dimensional (2D) Reynolds-averaged Navier–Stokes model [LANL's Besnard–Harlow–Rauenzahn (BHR) model] in order to describe an unsteady transitional Richtmyer–Meshkov (RM)-induced flow observed in on-going high-energy-density ensemble experiments performed on the OMEGA-EP facility. The experiments consist of a nominal single-mode perturbation (initial amplitude a 0 ≈ 10 and wavelength $λ$ = 100μm) with target-to-target variations in the surface roughness subjected to the RM instability with delayed Rayleigh–Taylor in a heavy-to-light configuration. Our strategy leverages high-resolution three-dimensional (3D) implicit large eddy simulations (ILES) simulations to initialize BHR-relevant parameters and subsequently validate the 2D BHR results against the 3D ILES simulations. A suite of five 3D ILES simulations corresponding to five experimental target profiles is undertaken to generate an ensemble dataset. Using ensemble averages from the 3D simulations to initialize the turbulent kinetic energy in the BHR model ( K 0 ) demonstrates the ability of the model to predict the time evolution of the interface as well as the density-specific-volume covariance, b . To quantify the sensitivity of the BHR results to the choice of K 0 and the initial turbulent length scale, S 0 , we execute a parameter sweep spanning four orders of magnitude for both S 0 and K 0 , generating a parameter space consisting of 26 simulations. The Pearson's correlation coefficient is used as a measure of discrepancy between the 2D BHR and 3D ILES simulations and reveals that the ranges 8≲S 0 ≲20 μm and 10 9 ≲K 0 ≲10 10 cm 2 /s 2 produce predictions that agree best with the 3D ILES results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A combined ensemble-volume average homogenization method for lattice structures with defects under dynamic and static loading

In the study of lattices structures, both experiments and numerical simulations are often conducted with small samples. Using combined ensemble and volume averaging, this work introduces a method to extract a macroscopic constitutive response of a lattice material from numerical simulations performed in periodic domains. The domain size needed to obtain statistically accurate results is investigated. Similar to molecular dynamics, the concept of the virial stress is introduced after homogenized equations are derived using the ensemble averaging method. Under static conditions, the virial stress is shown to agree with the volume averaged solid stress. Using the homogenization method, constitutive relations for this stress can be obtained from systems with uniform strains. Application of such obtained constitutive relations to more general cases results in an error proportional to the square of the ratio between the lattice length scale and the macroscopic length scale. Taking advantage of this property, numerical simulations are performed in systems with a uniform gradient of the average velocity. The volume average method is then used to accelerate convergence when studying lattices with defects. To avoid the artificial numerical time scale from the size of a representative volume element divided by the wave speed, a numerical scheme is developed to enforce a spatially uniform velocity gradient within the computational domain while allowing fluctuations of the velocity or displacement to develop naturally. To account for probability distribution of lattice defects, the stress is calculated as the ensemble-volume averaged value. For dynamic systems, energy dissipation properties are also studied.

36 MATERIALS SCIENCE↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

Linear Time-Invariant Models of a Large Cumulus Ensemble

Abstract Methods in system identification are used to obtain linear time-invariant state-space models that describe how horizontal averages of temperature and humidity of a large cumulus ensemble evolve with time under small forcing. The cumulus ensemble studied here is simulated with cloud-system-resolving models in radiative–convective equilibrium. The identified models extend steady-state linear response functions used in past studies and provide accurate descriptions of the transfer function, the noise model, and the behavior of cumulus convection when coupled with two-dimensional gravity waves. A novel procedure is developed to convert the state-space models into an interpretable form, which is used to elucidate and quantify memory in cumulus convection. The linear problem studied here serves as a useful reference point for more general efforts to obtain data-driven and interpretable parameterizations of cumulus convection.

Meteorology & Atmospheric Sciences↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

Non-asymptotic analysis of ensemble Kalman updates: effective dimension and localization

Many modern algorithms for inverse problems and data assimilation rely on ensemble Kalman updates to blend prior predictions with observed data. Ensemble Kalman methods often perform well with a small ensemble size, which is essential in applications where generating each particle is costly. This paper develops a non-asymptotic analysis of ensemble Kalman updates, which rigorously explains why a small ensemble size suffices if the prior covariance has moderate effective dimension due to fast spectrum decay or approximate sparsity. Here, we present our theory in a unified framework, comparing everal implementations of ensemble Kalman updates that use perturbed observations, square root filtering and localization. As part of our analysis, we develop new dimension-free covariance estimation bounds for approximately sparse matrices that may be of independent interest.

Mathematics↗

Overview of recent turbulence studies across multiple confinement modes at the ASDEX Upgrade tokamak using the Correlation Electron Cyclotron Emission diagnostic

This work presents an overview of recent and ongoing experimental measurements of core and edge turbulence across multiple confinement regimes using the Correlation Electron Cyclotron Emission (CECE) diagnostic at the ASDEX Upgrade (AUG) tokamak. A common goal among these investigations is to identify how the properties of the turbulent electron temperature fluctuations measured by CECE influence and regulate the unique transport characteristics of each confinement regime, including L-mode, I-mode, ELMy H-mode, and ELM-free H-mode. Optics and signal processing methods to aid in the analysis and interpretation of experimental turbulence results are also presented. These methods, and particularly the down-sampling and ensemble averaging method, are relevant to a wide variety of fusion and non-fusion applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

The WRF-Solar Ensemble Prediction System: Development, Test, and Validation

Providing reliable probabilistic solar radiation information is needed to improve management of the uncertainty and variability of solar generation. Thus, guidance on how to develop skillful and accurate ensemble forecasts is essential and it will ultimately contribute to integration of high amounts of solar energy on the grid. A team from the National Renewable Energy Laboratory and the National Center for Atmospheric Research had been collaborating to develop the WRF-Solar ensemble prediction system (WRF-Solar EPS) in the past three years to produce probabilistic solar irradiance forecasts and better predict solar energy by quantifying forecast uncertainty. The WRF-Solar EPS basically generates ensemble members for solar irradiance based on stochastic perturbations to provide intraday and day-ahead probabilistic forecasts. This study will present main research steps in developing the WRF-Solar EPS including: (a) tangent linear analysis for identifying key input variables of six WRF-Solar modules significantly related to predicting of cloud and solar irradiance, (b) combining stochastic perturbation technique with the WRF-Solar model, and (c) ensemble calibration method to decrease error and uncertainty of ensemble-based solar forecasts. The capability of WRF-Solar EPS is now updated to the most recent version of standard WRF model. This presentation will summarize comprehensive results from the evaluation of forecasts against the National Solar Radiation Data Base as well as ground-measured observations. Moreover, we will introduce the user's guide for WRF-Solar EPS (e.g., parameters to configure stochastic perturbations) and future extension of this research.

day-ahead forecast↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Early events in G-quadruplex folding captured by time-resolved small-angle X-ray scattering

Abstract Time-resolved small-angle X-ray experiments are reported here that capture and quantify a previously unknown rapid collapse of the unfolded oligonucleotide as an early step in the folding of hybrid 1 and hybrid 2 telomeric G-quadruplex structures. The rapid collapse, initiated by a pH jump, is characterized by an exponential decrease in the radius of gyration from 24.3 to 12.6 Å. The collapse is monophasic and is complete in <600 ms. Additional hand-mixing pH-jump kinetic studies show that slower kinetic steps follow the collapse. The folded and unfolded states at equilibrium were further characterized by SAXS studies and other biophysical tools, showing that G4 unfolding was complete at alkaline pH, but not in LiCl solution as is often claimed. The SAXS Ensemble Optimization Method analysis reveals models of the unfolded state as a dynamic ensemble of flexible oligonucleotide chains with a variety of transient hairpin structures. These results suggest a G4 folding pathway in which a rapid collapse, analogous to molten globule formation seen in proteins, is followed by a confined conformational search within the collapsed particle to form the native contacts ultimately found in the stable folded form.

Biochemistry & Molecular Biology↗

A Novel Data Segmentation Method for Data-driven Phase Identification

This paper presents a smart meter phase identification algorithm for two cases: meter-phase-label-known and meter-phase-label-unknown. To improve the identification accuracy, a data segmentation method is proposed to exclude data segments that are collected when the voltage correlation between smart meters on the same phase is weakened. Then, using the selected data segments, a hierarchical clustering method is used to calculate the correlation distances and cluster the smart meters. If the phase labels are unknown, a Connected-Triple-based Similarity (CTS) method is adapted to further improve the phase identification accuracy of the ensemble clustering method. The methods are developed and tested on both synthetic and real feeder data sets. Here, simulation results show that the proposed phase identification algorithm outperforms the state-of-the-art methods in both accuracy and robustness.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion↗

Analysis of a Computational Framework for Bayesian Inverse Problems: Ensemble Kalman Updates and MAP Estimators under Mesh Refinement

This paper analyzes a popular computational framework to solve infinite-dimensional Bayesian inverse problems, discretizing the prior and the forward model in a finite-dimensional weighted inner product space. We demonstrate the benefit of working on a weighted space by establishing operator-norm bounds for finite element and graph-based discretizations of Matérn-type priors and deconvolution forward models. For linear-Gaussian inverse problems, we develop a general theory to characterize the error in the approximation to the posterior. We also embed the computational framework into ensemble Kalman methods and MAP estimators for nonlinear inverse problems. Furthermore, our operator-norm bounds for prior discretizations guarantee the scalability and accuracy of these algorithms under mesh refinement.

Bayesian inverse problem↗