Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Application of a Split-Fiber Probe to Velocity Measurement in the NASA Research Compressor

A split-fiber probe was used to acquire unsteady data in a research compressor. The probe has two thin films deposited on a quartz cylinder 200 microns in diameter. A split-fiber probe allows simultaneous measurement of velocity magnitude and direction in a plane that is perpendicular to the sensing cylinder, because it has its circumference divided into two independent parts. Local heat transfer considerations indicated that the probe direction characteristic is linear in the range of flow incidence angles of +/- 35. Calibration tests confirmed this assumption. Of course, the velocity characteristic is nonlinear as is typical in thermal anemometry. The probe was used extensively in the NASA Glenn Research Center (GRC) low-speed, multistage axial compressor, and worked reliably during a test program of several months duration. The velocity and direction characteristics of the probe showed only minute changes during the entire test program. An algorithm was developed to decompose the probe signals into velocity magnitude and velocity direction. The averaged unsteady data were compared with data acquired by pneumatic probes. An overall excellent agreement between the averaged data acquired by a split-fiber probe and a pneumatic probe boosts confidence in the reliability of the unsteady content of the split-fiber probe data. To investigate the features of unsteady data, two methods were used: ensemble averaging and frequency analysis. The velocity distribution in a rotor blade passage was retrieved using the ensemble averaging method. Frequencies of excitation forces that may contribute to high cycle fatigue problems were identified by applying a fast Fourier transform to the absolute velocity data.

Lepicovsky, Jan

Device and Method for Gathering Ensemble Data Sets

An ensemble detector uses calibrated noise references to produce ensemble sets of data from which properties of non-stationary processes may be extracted. The ensemble detector comprising: a receiver; a switching device coupled to the receiver, the switching device configured to selectively connect each of a plurality of reference noise signals to the receiver; and a gain modulation circuit coupled to the receiver and configured to vary a gain of the receiver based on a forcing signal; whereby the switching device selectively connects each of the plurality of reference noise signals to the receiver to produce an output signal derived from the plurality of reference noise signals and the forcing signal.

Racette, Paul E.

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization

Analyzing Tropical Waves Using the Parallel Ensemble Empirical Model Decomposition Method: Preliminary Results from Hurricane Sandy

In this study, we discuss the performance of the parallel ensemble empirical mode decomposition (EMD) in the analysis of tropical waves that are associated with tropical cyclone (TC) formation. To efficiently analyze high-resolution, global, multiple-dimensional data sets, we first implement multilevel parallelism into the ensemble EMD (EEMD) and obtain a parallel speedup of 720 using 200 eight-core processors. We then apply the parallel EEMD (PEEMD) to extract the intrinsic mode functions (IMFs) from preselected data sets that represent (1) idealized tropical waves and (2) large-scale environmental flows associated with Hurricane Sandy (2012). Results indicate that the PEEMD is efficient and effective in revealing the major wave characteristics of the data, such as wavelengths and periods, by sifting out the dominant (wave) components. This approach has a potential for hurricane climate study by examining the statistical relationship between tropical waves and TC formation.

PEEMD

Ensemble Methodologies for Astronaut Cancer Risk Assessment in the face of Large Uncertainties

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

space radiation

Ensemble Cancer Risk Model for Astronaut Risk Assessment

A new approach to NASA space radiation risk modeling has successfully extended the current NASA probabilistic cancer risk model to an ensemble framework able to consider sub-model parameter uncertainty (e.g. uncertainty in a radiation quality parameter) as well as model-form uncertainty associated with differing theoretical or empirical formalisms (e.g. combined dose-rate and radiation quality effects). Ensemble methodologies are already widely used in weather prediction, modeling of infectious disease outbreaks, and certain terrestrial radiation protection applications to better understand how uncertainty may influence risk decision-making. Applying ensemble methodologies to space radiation risk projections offers the potential to efficiently incorporate emerging research results, allow for the incorporation of future (including international) models, improve uncertainty quantification for underlying sub-models developed against sparse experimental data, and reduce the impact of subjective bias on risk projections. Moreover, risk forecasting across an ensemble of multiple predictive models can provide stakeholders additional information on risk acceptance if current health/medical standards cannot be met or the level of knowledge doesn’t permit a specific risk or exposure limit to be developed for future space exploration missions. In this work, ensemble risk projections implementing multiple sub-models of radiation quality, dose and dose-rate effectiveness factors, excess risk, and latency as ensemble members are presented. Initial consensus methods for ensemble model weights and correlations to account for individual model bias are discussed. In these analyses, the ensemble forecast compares well to results from NASA's current operational cancer risk projection model used to assess permissible exposure limits and permissible mission durations for astronauts. However, a large range of projected risk values are obtained at the upper 95th confidence level where models must extrapolate beyond available biological data sets; closer agreement is seen at the median + one sigma due to the inherent similarities in available models. Future work, including the addition of new models and methods for statistical correlation between predictive members are discussed to define alternate ways of thinking about risk and ‘acceptable’ uncertainty with respect to NASA’s current permissible exposure limits.

Lisa C Simonsen

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods

An interplanetary magnetic field ensemble at 1 AU

A method for calculation ensemble averages from magnetic field data is described. A data set comprising approximately 16 months of nearly continuous ISEE-3 magnetic field data is used in this study. Individual subintervals of this data, ranging from 15 hours to 15.6 days comprise the ensemble. The sole condition for including each subinterval in the averages is the degree to which it represents a weakly time-stationary process. Averages obtained by this method are appropriate for a turbulence description of the interplanetary medium. The ensemble average correlation length obtained from all subintervals is found to be 4.9 x 10 to the 11th cm. The average value of the variances of the magnetic field components are in the approximate ratio 8:9:10, where the third component is the local mean field direction. The correlation lengths and variances are found to have a systematic variation with subinterval duration, reflecting the important role of low-frequency fluctuations in the interplanetary medium.

Matthaeus, W. H.

An interplanetary magnetic field ensemble at 1 AU

A method for calculation ensemble averages from magnetic field data is described. A data set comprising approximately 16 months of nearly continuous ISEE-3 magnetic field data is used in this study. Individual subintervals of this data, ranging from 15 hours to 15.6 days comprise the ensemble. The sole condition for including each subinterval in the averages is the degree to shich it represents a weakly time-stationary process. Averages obtained by this method are appropriate for a turbulence description of the interplanetary medium. The ensemble average correlation length obtained from all subintervals is found to be 4.9 x 10 to the 11th cm. The average value of the variances of the magnetic field components are in the approximate ratio 8:9:10, where the third component is the local mean field direction. The correlation lengths and variances are found to have a systematic variation with subinterval duration, reflecting the important role of low-frequency fluctuations in the interplanetary medium.

Matthaeus, W. H.

Overview of recent turbulence studies across multiple confinement modes at the ASDEX Upgrade tokamak using the Correlation Electron Cyclotron Emission diagnostic

This work presents an overview of recent and ongoing experimental measurements of core and edge turbulence across multiple confinement regimes using the Correlation Electron Cyclotron Emission (CECE) diagnostic at the ASDEX Upgrade (AUG) tokamak. A common goal among these investigations is to identify how the properties of the turbulent electron temperature fluctuations measured by CECE influence and regulate the unique transport characteristics of each confinement regime, including L-mode, I-mode, ELMy H-mode, and ELM-free H-mode. Optics and signal processing methods to aid in the analysis and interpretation of experimental turbulence results are also presented. These methods, and particularly the down-sampling and ensemble averaging method, are relevant to a wide variety of fusion and non-fusion applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions

Early events in G-quadruplex folding captured by time-resolved small-angle X-ray scattering

Abstract Time-resolved small-angle X-ray experiments are reported here that capture and quantify a previously unknown rapid collapse of the unfolded oligonucleotide as an early step in the folding of hybrid 1 and hybrid 2 telomeric G-quadruplex structures. The rapid collapse, initiated by a pH jump, is characterized by an exponential decrease in the radius of gyration from 24.3 to 12.6 Å. The collapse is monophasic and is complete in <600 ms. Additional hand-mixing pH-jump kinetic studies show that slower kinetic steps follow the collapse. The folded and unfolded states at equilibrium were further characterized by SAXS studies and other biophysical tools, showing that G4 unfolding was complete at alkaline pH, but not in LiCl solution as is often claimed. The SAXS Ensemble Optimization Method analysis reveals models of the unfolded state as a dynamic ensemble of flexible oligonucleotide chains with a variety of transient hairpin structures. These results suggest a G4 folding pathway in which a rapid collapse, analogous to molten globule formation seen in proteins, is followed by a confined conformational search within the collapsed particle to form the native contacts ultimately found in the stable folded form.

Biochemistry & Molecular Biology

Effects of a Rotating Aerodynamic Probe on the Flow Field of a Compressor Rotor

An investigation of distortions of the rotor exit flow field caused by an aerodynamic probe mounted in the rotor is described in this paper. A rotor total pressure Kiel probe, mounted on the rotor hub and extending up to the mid-span radius of a rotor blade channel, generates a wake that forms additional flow blockage. Three types of high-response aerodynamic probes were used to investigate the distorted flow field behind the rotor. These probes were: a split-fiber thermo-anemometric probe to measure velocity and flow direction, a total pressure probe, and a disk probe for in-flow static pressure measurement. The signals acquired from these high-response probes were reduced using an ensemble averaging method based on a once per rotor revolution signal. The rotor ensemble averages were combined to construct contour plots for each rotor channel of the rotor tested. In order to quantify the rotor probe effects, the contour plots for each individual rotor blade passage were averaged into a single value. The distribution of these average values along the rotor circumference is a measure of changes in the rotor exit flow field due to the presence of a probe in the rotor. These distributions were generated for axial flow velocity and for static pressure.

Lepicovsky, Jan

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion