Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A combined ensemble-volume average homogenization method for lattice structures with defects under dynamic and static loading

In the study of lattices structures, both experiments and numerical simulations are often conducted with small samples. Using combined ensemble and volume averaging, this work introduces a method to extract a macroscopic constitutive response of a lattice material from numerical simulations performed in periodic domains. The domain size needed to obtain statistically accurate results is investigated. Similar to molecular dynamics, the concept of the virial stress is introduced after homogenized equations are derived using the ensemble averaging method. Under static conditions, the virial stress is shown to agree with the volume averaged solid stress. Using the homogenization method, constitutive relations for this stress can be obtained from systems with uniform strains. Application of such obtained constitutive relations to more general cases results in an error proportional to the square of the ratio between the lattice length scale and the macroscopic length scale. Taking advantage of this property, numerical simulations are performed in systems with a uniform gradient of the average velocity. The volume average method is then used to accelerate convergence when studying lattices with defects. To avoid the artificial numerical time scale from the size of a representative volume element divided by the wave speed, a numerical scheme is developed to enforce a spatially uniform velocity gradient within the computational domain while allowing fluctuations of the velocity or displacement to develop naturally. To account for probability distribution of lattice defects, the stress is calculated as the ensemble-volume averaged value. For dynamic systems, energy dissipation properties are also studied.

36 MATERIALS SCIENCE↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

Linear Time-Invariant Models of a Large Cumulus Ensemble

Abstract Methods in system identification are used to obtain linear time-invariant state-space models that describe how horizontal averages of temperature and humidity of a large cumulus ensemble evolve with time under small forcing. The cumulus ensemble studied here is simulated with cloud-system-resolving models in radiative–convective equilibrium. The identified models extend steady-state linear response functions used in past studies and provide accurate descriptions of the transfer function, the noise model, and the behavior of cumulus convection when coupled with two-dimensional gravity waves. A novel procedure is developed to convert the state-space models into an interpretable form, which is used to elucidate and quantify memory in cumulus convection. The linear problem studied here serves as a useful reference point for more general efforts to obtain data-driven and interpretable parameterizations of cumulus convection.

Meteorology & Atmospheric Sciences↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

Analyzing Tropical Waves Using the Parallel Ensemble Empirical Model Decomposition Method: Preliminary Results from Hurricane Sandy

In this study, we discuss the performance of the parallel ensemble empirical mode decomposition (EMD) in the analysis of tropical waves that are associated with tropical cyclone (TC) formation. To efficiently analyze high-resolution, global, multiple-dimensional data sets, we first implement multilevel parallelism into the ensemble EMD (EEMD) and obtain a parallel speedup of 720 using 200 eight-core processors. We then apply the parallel EEMD (PEEMD) to extract the intrinsic mode functions (IMFs) from preselected data sets that represent (1) idealized tropical waves and (2) large-scale environmental flows associated with Hurricane Sandy (2012). Results indicate that the PEEMD is efficient and effective in revealing the major wave characteristics of the data, such as wavelengths and periods, by sifting out the dominant (wave) components. This approach has a potential for hurricane climate study by examining the statistical relationship between tropical waves and TC formation.

PEEMD↗

Ensemble Methodologies for Astronaut Cancer Risk Assessment in the face of Large Uncertainties

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

space radiation↗

Ensemble Cancer Risk Model for Astronaut Risk Assessment

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

Lisa C Simonsen↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

An interplanetary magnetic field ensemble at 1 AU

A method for calculation ensemble averages from magnetic field data is described. A data set comprising approximately 16 months of nearly continuous ISEE-3 magnetic field data is used in this study. Individual subintervals of this data, ranging from 15 hours to 15.6 days comprise the ensemble. The sole condition for including each subinterval in the averages is the degree to which it represents a weakly time-stationary process. Averages obtained by this method are appropriate for a turbulence description of the interplanetary medium. The ensemble average correlation length obtained from all subintervals is found to be 4.9 x 10 to the 11th cm. The average value of the variances of the magnetic field components are in the approximate ratio 8:9:10, where the third component is the local mean field direction. The correlation lengths and variances are found to have a systematic variation with subinterval duration, reflecting the important role of low-frequency fluctuations in the interplanetary medium.

Matthaeus, W. H.↗

An interplanetary magnetic field ensemble at 1 AU

A method for calculation ensemble averages from magnetic field data is described. A data set comprising approximately 16 months of nearly continuous ISEE-3 magnetic field data is used in this study. Individual subintervals of this data, ranging from 15 hours to 15.6 days comprise the ensemble. The sole condition for including each subinterval in the averages is the degree to shich it represents a weakly time-stationary process. Averages obtained by this method are appropriate for a turbulence description of the interplanetary medium. The ensemble average correlation length obtained from all subintervals is found to be 4.9 x 10 to the 11th cm. The average value of the variances of the magnetic field components are in the approximate ratio 8:9:10, where the third component is the local mean field direction. The correlation lengths and variances are found to have a systematic variation with subinterval duration, reflecting the important role of low-frequency fluctuations in the interplanetary medium.

Matthaeus, W. H.↗

Non-asymptotic analysis of ensemble Kalman updates: effective dimension and localization

Many modern algorithms for inverse problems and data assimilation rely on ensemble Kalman updates to blend prior predictions with observed data. Ensemble Kalman methods often perform well with a small ensemble size, which is essential in applications where generating each particle is costly. This paper develops a non-asymptotic analysis of ensemble Kalman updates, which rigorously explains why a small ensemble size suffices if the prior covariance has moderate effective dimension due to fast spectrum decay or approximate sparsity. Here, we present our theory in a unified framework, comparing everal implementations of ensemble Kalman updates that use perturbed observations, square root filtering and localization. As part of our analysis, we develop new dimension-free covariance estimation bounds for approximately sparse matrices that may be of independent interest.

Mathematics↗

Overview of recent turbulence studies across multiple confinement modes at the ASDEX Upgrade tokamak using the Correlation Electron Cyclotron Emission diagnostic

This work presents an overview of recent and ongoing experimental measurements of core and edge turbulence across multiple confinement regimes using the Correlation Electron Cyclotron Emission (CECE) diagnostic at the ASDEX Upgrade (AUG) tokamak. A common goal among these investigations is to identify how the properties of the turbulent electron temperature fluctuations measured by CECE influence and regulate the unique transport characteristics of each confinement regime, including L-mode, I-mode, ELMy H-mode, and ELM-free H-mode. Optics and signal processing methods to aid in the analysis and interpretation of experimental turbulence results are also presented. These methods, and particularly the down-sampling and ensemble averaging method, are relevant to a wide variety of fusion and non-fusion applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

The WRF-Solar Ensemble Prediction System: Development, Test, and Validation

Providing reliable probabilistic solar radiation information is needed to improve management of the uncertainty and variability of solar generation. Thus, guidance on how to develop skillful and accurate ensemble forecasts is essential and it will ultimately contribute to integration of high amounts of solar energy on the grid. A team from the National Renewable Energy Laboratory and the National Center for Atmospheric Research had been collaborating to develop the WRF-Solar ensemble prediction system (WRF-Solar EPS) in the past three years to produce probabilistic solar irradiance forecasts and better predict solar energy by quantifying forecast uncertainty. The WRF-Solar EPS basically generates ensemble members for solar irradiance based on stochastic perturbations to provide intraday and day-ahead probabilistic forecasts. This study will present main research steps in developing the WRF-Solar EPS including: (a) tangent linear analysis for identifying key input variables of six WRF-Solar modules significantly related to predicting of cloud and solar irradiance, (b) combining stochastic perturbation technique with the WRF-Solar model, and (c) ensemble calibration method to decrease error and uncertainty of ensemble-based solar forecasts. The capability of WRF-Solar EPS is now updated to the most recent version of standard WRF model. This presentation will summarize comprehensive results from the evaluation of forecasts against the National Solar Radiation Data Base as well as ground-measured observations. Moreover, we will introduce the user's guide for WRF-Solar EPS (e.g., parameters to configure stochastic perturbations) and future extension of this research.

day-ahead forecast↗