Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimizing ensemble NV − spin properties of fluorescent diamond microparticles by systematic low pressure high temperature annealing

Low pressure high temperature annealing is a means for driving nitrogen and defect diffusion in diamond to reduce internal lattice damage without the need for technically complicated high-pressure cells. Herein, we perform a systematic time (5, 15, and 30 min) and temperature (1200 °C–1800 °C) study of effects of low-pressure high temperature annealing on photoluminescence, spin concentrations, and spin relaxation properties of NV centers in ca. 3 μm synthetic type 1b diamond particles. Annealing in the temperature range of ca. 1400 °C–1700 °C for even 5 min leads to a higher optically detected magnetic resonance contrast as compared to standard annealing at 900 °C for 2 h. Particles annealed at 1700 °C for 5 min exhibit a contrast close to about 13% as compared to about 9% for those annealed at 900 °C for 2 h. A reduction in the zero-field splitting strain parameter from E ≈ 4.5 MHz to ≈ 2.5 MHz and spectral linewidth from Δν ≈ 7 MHz to ≈ 4 MHz are observed even after 5 min annealing at 1700 °C. Improvements in these spectral parameters resulted in a roughly 2-fold reduction in the noise level of temperature monitoring experiment utilizing an ensemble of NV centers in the particles. Annealing in the temperature range of 1600 °C for 15 or 30 min or 1700 °C for 5 min resulted in NV T 1 relaxation times approaching ca. 5 ms typically observed for bulk diamond. Quantitative electron paramagnetic resonance (EPR) allowed for estimations of thermal activation energies of paramagnetic center annihilation. Monitoring the primary defect concentration (P1 and other defects with half integer spins) and utilizing second order kinetic modeling, an activation energy of 3.63 ± 0.28 eV was estimated. Alternatively, using the NV half field EPR signal and first order kinetic modeling, a similar activation energy 3.89 ± 0.29 eV was estimated.

NV centers↗

Designing an Optimal Ensemble Strategy for GMAO S2S Forecast System

The NASA Global Modeling and Assimilation Office (GMAO) Sub-seasonal to Seasonal (S2S) prediction system is being readied for a major upgrade. An important factor in successful extended range forecasting is the definition of the ensemble. Our overall strategy is to run a relatively large ensemble of about 40 members up to 3 months (focusing on the sub-seasonal forecast problem), after which we sub-sample the ensemble, and continue the forecast with about 10 members (up to 12 months). Here we present the results of our testing of various ways to generate the initial perturbations and the validation of a stratified sampling approach for choosing the members of the smaller ensemble. For the initialization of the ensemble we propose a combination of lagged and burst initial conditions. To generate perturbations for the burst ensemble members we used scaled differences of pairs of analysis states (chosen randomly from the corresponding season) separated by 1-10 days. We consider perturbing separately the atmosphere and the ocean, or both. By varying the separation times between the analysis states, we are able to produce perturbations that resemble well-known modes of variability. Focusing on the ENSO SST indices, we found that all types of perturbations are important for the ensemble spread with, however, considerable differences in the timing of the impacts on spread for the atmospheric and oceanic perturbations.Our initial (larger) ensemble size was determined so as to maximize the skill of predicting some of the leading modes of boreal winter atmospheric modes (namely the NAO, PNA and AO). Since it is not feasible for us to run with the larger ensemble beyond about 3 months, we employ a stratified sampling procedure that identifies the emerging directions of error growth to subset the ensemble. By comparing the results from the stratified ensemble with that of the randomly sampled ensemble of the same size, we find that the former provides substantially better estimates the mean of the original large ensemble.

Borovikov, Anna↗

Designing an Optimal Ensemble Strategy for GMAO S2S Forecast System

GMAO Sub/Seasonal prediction system (S2S) is being readied for a major upgrade to GEOS-S2S Version 3. An important factor in successful extended range forecast is the definition of an ensemble For initialization of the ensemble we propose a combination of lagged and burst initial conditions. We plan to run a relatively large ensemble of 40 members for sub-seasonal forecast (up to 3 months), at which point we sub-sample the ensemble, and continue the forecast with 10 members (up to 12 months). Here we present the results of the extensive testing of various ways to generate the perturbations to the initial conditions and the validation of the stratified sampling strategy we chose.To generate perturbations for the burst ensemble members we used scaled differences of pairs of analysis states separated by 1-10 days, randomly chosen from a corresponding season. We considered perturbing separately only the atmospheric fields or only the ocean or both of the forecast initial conditions. Considering varying separation times between the analysis states, we were able to produce perturbations sampling various modes of variability. Focusing on the ENSO SST indices, we found that all types of perturbations are important for the ensemble spread.Our ensemble size for sub-seasonal forecasts was determined as to maximize the skill of predicting some of the leading modes of boreal winter atmospheric modes, NAO, PNA and AO. It is not feasible to run equally large ensemble for seasonal forecasts. Using a stratified sampling procedure we can identify the emerging directions of error growth. By comparing the stratified ensemble with randomly sampled ensemble of the same size, we were able to show that the former better estimates the mean of the original large ensemble.

Borovikov, Anna↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

Error Estimation of An Ensemble Statistical Seasonal Precipitation Prediction Model

This NASA Technical Memorandum describes an optimal ensemble canonical correlation forecasting model for seasonal precipitation. Each individual forecast is based on the canonical correlation analysis (CCA) in the spectral spaces whose bases are empirical orthogonal functions (EOF). The optimal weights in the ensemble forecasting crucially depend on the mean square error of each individual forecast. An estimate of the mean square error of a CCA prediction is made also using the spectral method. The error is decomposed onto EOFs of the predictand and decreases linearly according to the correlation between the predictor and predictand. Since new CCA scheme is derived for continuous fields of predictor and predictand, an area-factor is automatically included. Thus our model is an improvement of the spectral CCA scheme of Barnett and Preisendorfer. The improvements include (1) the use of area-factor, (2) the estimation of prediction error, and (3) the optimal ensemble of multiple forecasts. The new CCA model is applied to the seasonal forecasting of the United States (US) precipitation field. The predictor is the sea surface temperature (SST). The US Climate Prediction Center's reconstructed SST is used as the predictor's historical data. The US National Center for Environmental Prediction's optimally interpolated precipitation (1951-2000) is used as the predictand's historical data. Our forecast experiments show that the new ensemble canonical correlation scheme renders a reasonable forecasting skill. For example, when using September-October-November SST to predict the next season December-January-February precipitation, the spatial pattern correlation between the observed and predicted are positive in 46 years among the 50 years of experiments. The positive correlations are close to or greater than 0.4 in 29 years, which indicates excellent performance of the forecasting model. The forecasting skill can be further enhanced when several predictors are used.

Shen, Samuel S. P.↗

A Canonical Ensemble Correlation Prediction Model for Seasonal Precipitation Anomaly

This report describes an optimal ensemble forecasting model for seasonal precipitation and its error estimation. Each individual forecast is based on the canonical correlation analysis (CCA) in the spectral spaces whose bases are empirical orthogonal functions (EOF). The optimal weights in the ensemble forecasting crucially depend on the mean square error of each individual forecast. An estimate of the mean square error of a CCA prediction is made also using the spectral method. The error is decomposed onto EOFs of the predictand and decreases linearly according to the correlation between the predictor and predictand. This new CCA model includes the following features: (1) the use of area-factor, (2) the estimation of prediction error, and (3) the optimal ensemble of multiple forecasts. The new CCA model is applied to the seasonal forecasting of the United States precipitation field. The predictor is the sea surface temperature.

Shen, Samuel S. P.↗

A New Ensemble Canonical Correlation Prediction Scheme for Seasonal Precipitation

Department of Mathematical Sciences, University of Alberta, Edmonton, Canada This paper describes the fundamental theory of the ensemble canonical correlation (ECC) algorithm for the seasonal climate forecasting. The algorithm is a statistical regression sch eme based on maximal correlation between the predictor and predictand. The prediction error is estimated by a spectral method using the basis of empirical orthogonal functions. The ECC algorithm treats the predictors and predictands as continuous fields and is an improvement from the traditional canonical correlation prediction. The improvements include the use of area-factor, estimation of prediction error, and the optimal ensemble of multiple forecasts. The ECC is applied to the seasonal forecasting over various parts of the world. The example presented here is for the North America precipitation. The predictor is the sea surface temperature (SST) from different ocean basins. The Climate Prediction Center's reconstructed SST (1951-1999) is used as the predictor's historical data. The optimally interpolated global monthly precipitation is used as the predictand?s historical data. Our forecast experiments show that the ECC algorithm renders very high skill and the optimal ensemble is very important to the high value.

Kim, Kyu-Myong↗

Third international challenge to model the medium- to long-range transport of radioxenon to four Comprehensive Nuclear-Test-Ban Treaty monitoring stations

In 2015 and 2016, atmospheric transport modeling challenges were conducted in the context of the Comprehensive Nuclear-Test-Ban Treaty (CTBT) verification, however, with a more limited scope with respect to emission inventories, simulation period and number of relevant samples (i.e., those above the Minimum Detectable Concentration (MDC)) involved. Therefore, a more comprehensive atmospheric transport modeling challenge was organized in 2019. Stack release data of Xe-133 were provided by the Institut National des Radioéléments/IRE (Belgium) and the Canadian Nuclear Laboratories/CNL (Canada) and accounted for in the simulations over a three (mandatory) or six (optional) months period. Best estimate emissions of additional facilities (radiopharmaceutical production and nuclear research facilities, commercial reactors or relevant research reactors) of the Northern Hemisphere were included as well. Model results were compared with observed atmospheric activity concentrations at four International Monitoring System (IMS) stations located in Europe and North America with overall considerable influence of IRE and/or CNL emissions for evaluation of the participants’ runs. Participants were prompted to work with controlled and harmonized model set-ups to make runs more comparable, but also to increase diversity. It was found that using the stack emissions of IRE and CNL with daily resolution does not lead to better results than disaggregating annual emissions of these two facilities taken from the literature if an overall score for all stations covering all valid observed samples is considered. A moderate benefit of roughly 10% is visible in statistical scores for samples influenced by IRE and/or CNL to at least 50% and there can be considerable benefit for individual samples. Effects of transport errors, not properly characterized remaining emitters and long IMS sampling times (12–24 h) undoubtedly are in contrast to and reduce the benefit of high-quality IRE and CNL stack data. Complementary best estimates for remaining emitters push the scores up by 18% compared to just considering IRE and CNL emissions alone. Despite the efforts undertaken the full multi-model ensemble built is highly redundant. An ensemble based on a few arbitrary runs is sufficient to model the Xe-133 background at the stations investigated. The effective ensemble size is below five. An optimized ensemble at each station has on average slightly higher skill compared to the full ensemble. However, the improvement (maximum of 20% and minimum of 3% in RMSE) in skill is likely being too small for being exploited for an independent period.

54 ENVIRONMENTAL SCIENCES↗

The optimization of model ensemble composition and size can enhance the robustness of crop yield projections

Linked climate and crop simulation models are widely used to assess the impact of climate change on agriculture. However, it is unclear how ensemble configurations (model composition and size) influence crop yield projections and uncertainty. Here, we investigate the influences of ensemble configurations on crop yield projections and modeling uncertainty from Global Gridded Crop Models and Global Climate Models under future climate change. We performed a cluster analysis to identify distinct groups of ensemble members based on their projected outcomes, revealing unique patterns in crop yield projections and corresponding uncertainty levels, particularly for wheat and soybean. Furthermore, our findings suggest that approximately six Global Gridded Crop Models and 10 Global Climate Models are sufficient to capture modeling uncertainty, while a cluster-based selection of 3-4 Global Gridded Crop Models effectively represents the full ensemble. The contribution of individual Global Gridded Crop Models to overall uncertainty varies depending on region and crop type, emphasizing the importance of considering the impact of specific models when selecting models for local-scale applications. Our results emphasize the importance of model composition and ensemble size in identifying the primary sources of uncertainty in crop yield projections, offering valuable guidance for optimizing ensemble configurations in climate-crop modeling studies tailored to specific applications.

Agriculture↗

Early events in G-quadruplex folding captured by time-resolved small-angle X-ray scattering

Abstract Time-resolved small-angle X-ray experiments are reported here that capture and quantify a previously unknown rapid collapse of the unfolded oligonucleotide as an early step in the folding of hybrid 1 and hybrid 2 telomeric G-quadruplex structures. The rapid collapse, initiated by a pH jump, is characterized by an exponential decrease in the radius of gyration from 24.3 to 12.6 Å. The collapse is monophasic and is complete in <600 ms. Additional hand-mixing pH-jump kinetic studies show that slower kinetic steps follow the collapse. The folded and unfolded states at equilibrium were further characterized by SAXS studies and other biophysical tools, showing that G4 unfolding was complete at alkaline pH, but not in LiCl solution as is often claimed. The SAXS Ensemble Optimization Method analysis reveals models of the unfolded state as a dynamic ensemble of flexible oligonucleotide chains with a variety of transient hairpin structures. These results suggest a G4 folding pathway in which a rapid collapse, analogous to molten globule formation seen in proteins, is followed by a confined conformational search within the collapsed particle to form the native contacts ultimately found in the stable folded form.

Biochemistry & Molecular Biology↗

Protein folding from heterogeneous unfolded state revealed by time-resolved X-ray solution scattering

One of the most challenging tasks in biological science is to understand how a protein folds. In theoretical studies, the hypothesis adopting a funnel-like free-energy landscape has been recognized as a prominent scheme for explaining protein folding in views of both internal energy and conformational heterogeneity of a protein. Despite numerous experimental efforts, however, comprehensively studying protein folding with respect to its global conformational changes in conjunction with the heterogeneity has been elusive. Here we investigate the redox-coupled folding dynamics of equine heart cytochrome c (cyt-c) induced by external electron injection by using time-resolved X-ray solution scattering. A systematic kinetic analysis unveils a kinetic model for its folding with a stretched exponential behavior during the transition toward the folded state. With the aid of the ensemble optimization method combined with molecular dynamics simulations, we found that during the folding the heterogeneously populated ensemble of the unfolded state is converted to a narrowly populated ensemble of folded conformations. These observations obtained from the kinetic and the structural analyses of X-ray scattering data reveal that the folding dynamics of cyt-c accompanies many parallel pathways associated with the heterogeneously populated ensemble of unfolded conformations, resulting in the stretched exponential kinetics at room temperature. This finding provides direct evidence with a view to microscopic protein conformations that the cyt-c folding initiates from a highly heterogeneous unfolded state, passes through still diverse intermediate structures, and reaches structural homogeneity by arriving at the folded state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Global fire modelling and control attributions based on the ensemble machine learning and satellite observations

Contemporary fire dynamics is one of the most complex and least understood land surface phenomena. Global fire controls related to climate, vegetation, and anthropogenic activity are usually intertwined, and difficult to disentangle in a quantitative way. Here, we leveraged an ensemble of five machine learning (ML) models and multiple satellite-based observations to conduct global fire modeling for three fire metrics (burned area, fire number, and fire size), and quantified driving mechanisms underlying annual fire changes in a spatially resolved manner for the period 2003–2019. Ensemble learning is a meta-approach that combines multiple ML predictions to improve accuracy, robustness, and generalization performance. We found that the optimized ensemble ML well reproduced annual dynamics of global burned area (R 2 = 0.90, P < 0.001), total fire numbers (R 2 = 0.86, P < 0.001), and averaged fire size (R 2 = 0.70, P < 0.001). Additionally, the ensemble ML captured key spatial patterns of multi-year mean magnitudes, annual variabilities, anomalies, and trends for different fire metrics. Our ML-based fire attributions further highlighted the dominant role of enhanced anthropogenic activity in reducing global burned area (–1.9 Mha/yr, P < 0.01), followed by climate control (–1.3 Mha/yr, P < 0.01) and insignificant positive vegetation control (0.4 Mha/yr, P = 0.60). Spatially, climate dominated a much larger burned area (53.7%) than human (23.4%) or vegetation control (22.9%); however, the counteracting effects from regional wetting and drying trends weakened the net climate impacts on global burned area. The fire number and fire size exhibited similar spatial control patterns with burned area; globally, however, fire number tended to be more affected by climate while fire size more influenced by human activities. Overall, our study confirmed the feasibility and efficiency of ensemble ML in global fire modeling and subsequent control attributions, providing a better understanding of contemporary fire regimes and contributing to robust fire projections in a changing environment.

54 ENVIRONMENTAL SCIENCES↗

Avatar Tools

Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Siefert, Christopher↗

Decadal Prediction Skill in the GEOS-5 Forecast System

A suite of decadal predictions has been conducted with the NASA Global Modeling and Assimilation Office's (GMAO's) GEOS-5 Atmosphere-Ocean general circulation model. The hind casts are initialized every December 1st from 1959 to 2010, following the CMIP5 experimental protocol for decadal predictions. The initial conditions are from a multivariate ensemble optimal interpolation ocean and sea-ice reanalysis, and from GMAO's atmospheric reanalysis, the modern-era retrospective analysis for research and applications. The mean forecast skill of a three-member-ensemble is compared to that of an experiment without initialization but also forced with observed greenhouse gases. The results show that initialization increases the forecast skill of North Atlantic sea surface temperature compared to the uninitialized runs, with the increase in skill maintained for almost a decade over the subtropical and mid-latitude Atlantic. On the other hand, the initialization reduces the skill in predicting the warming trend over some regions outside the Atlantic. The annual-mean Atlantic meridional overturning circulation index, which is defined here as the maximum of the zonally-integrated overturning stream function at mid-latitude, is predictable up to a 4-year lead time, consistent with the predictable signal in upper ocean heat content over the North Atlantic. While the 6- to 9-year forecast skill measured by mean squared skill score shows 50 percent improvement in the upper ocean heat content over the subtropical and mid-latitude Atlantic, prediction skill is relatively low in the sub-polar gyre. This low skill is due in part to features in the spatial pattern of the dominant simulated decadal mode in upper ocean heat content over this region that differ from observations. An analysis of the large-scale temperature budget shows that this is the result of a model bias, implying that realistic simulation of the climatological fields is crucial for skillful decadal forecasts.

Decadal Prediction↗