Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Quantifying the Known Unknown: Including Marine Sources of Greenhouse Gases in Climate Modeling

Researchers have recently estimated that Arctic submarine permafrost currently traps 60 billion tons of methane and contains 560 billion tons of organic carbon in seafloor sediments and soil, a giant pool of carbon with potentially large feedbacks on the climate system. Unlike terrestrial permafrost, the submarine permafrost system has remained a “known unknown” because of the difficulty in acquiring samples and measurements. Consequently, this potentially large carbon stock never yet considered in global climate models or policy discussions, represents a real wildcard in our understanding of Earth’s climate. This report summarizes our group’s effort at developing a numerical modeling framework designed to produce a first-of-its-kind estimate of Arctic methane gas releases from the marine sediments to the water column, and potentially to the atmosphere, where positive climate feedback may occur. Newly developed modeling capability supported by the Laboratory Directed Research and Development (LDRD) program at Sandia National Laboratories now gives us the ability to probabilistically map gas distribution and quantity in the seabed by using a hybrid approach of geospatial machine learning, and predictive numerical thermodynamic ensemble modeling. The novelty in this approach is its ability to produce maps of useful data in regions that are only sparsely sampled, a common challenge in the Arctic, and a major obstacle to progress in the past. By applying this model to the circum-Arctic continental shelves and integrating the flux of free gas from in situ methanogenesis and dissociating gas hydrates from the sediment column under climate forcing, we can provide the most reliable estimate of a spatially and temporally varying source term for greenhouse gas flux that can be used by global oceanographic circulation and Earth system models (such as DOE’s E3SM). The result will allow us to finally tackle the wildcard of the submarine permafrost carbon system, and better inform us about the severity of future national security threats that sustained climate change poses.

54 ENVIRONMENTAL SCIENCES↗

Experimental investigation of self-induced transparency and pulse delay in ruby.

We have investigated the self-induced transparency effect in ruby over a range of input energies which range from linear absorption to full transparency. The transmission, pulse delay, and pulse broadening were studied as a function of input energy. The transition region is narrower than that found in similar studies of the CO2/SF6 system; this is consistent with predictions based on ensembles of two-level systems. Included are the first pulse-delay and pulse-broadening curves to be obtained for the ruby system.

Asher, I. M.↗

Performance of Soviet and US hydrogen masers

The frequencies of Soviet- and U.S.-built hydrogen masers located at the Smithsonian Astrophysical Observatory and at the United States Naval Observatory (USNO) were compared with each other and, via Global Positioning System (GPS) common-view measurements, with three primary frequency-reference scales. The best masers were found to have fractional frequency stabilities as low as 6 times 10(exp -16) for averaging times of approximately 10(exp 4) s. Members of the USNO maser ensemble provided frequency prediction better than 1 times 10(exp 14) for periods up to a few weeks. The frequency residuals of these masers, after removal of frequency drift and rate of change of drift, had stabilities of a few parts in 10(exp -15), with serveral masers achieving residual stabilities well below 1 times 10(exp -15) for intervals from 10(exp 5)s to 2 times 10(exp 6)s. The fractional frequency drifts of the 13 masers studied, relative to the primary reference standards, ranged from -0.2 times 10(exp -15)/day to +9.6 times 10(exp -15)/day.

Uljanov, Adolph A.↗

The evolution of voids in the adhesion approximation

We apply the adhesion approximation to study the formation and evolution of voids in the universe. Our simulations-carried out using 128(exp 3) particles in a cubical box with side 128 Mpc-indicate that the void spectrum evolves with time and that the mean void size in the standard Cosmic Background Explorer Satellite (COBE)-normalized cold dark matter (CDM) model with H(sub 50) = 1 scals approximately as bar D(z) = bar D(sub zero)/(1+2)(exp 1/2), where bar D(sub zero) approximately = 10.5 Mpc. Interestingly, we find a strong correlation between the sizes of voids and the value of the primordial gravitational potential at void centers. This observation could in principle, pave the way toward reconstructing the form of the primordialpotential from a knowledge of the observed void spectrum. Studying the void spectrum at different cosmological epochs, for spectra with a built in k-space cutoff we find that the number of voids in a representative volume evolves with time. The mean number of voids first increases until a maximum value is reached (indicating that the formation of cellular structure is complete), and then begins to decrease as clumps and filaments erge leading to hierarchical clustering and the subsequent elimination of small voids. The cosmological epoch characterizing the completion of cellular structure occurs when the length scale going nonlinear approaches the mean distance between peaks of the gravitaional potential. A central result of this paper is that voids can be populated by substructure such as mini-sheets and filaments, which run through voids. The number of such mini-pancakes that pass through a given void can be measured by the genus characteristic of an individual void which is an indicator of the topology of a given void in intial (Lagrangian) space. Large voids have on an average a larger measure than smaller voids indicating more substructure within larger voids relative to smaller ones. We find that the topology of individual voids is strongly epoch dependent, with void topologies generally simplifying with time. This means that as voids grow older they become progressively more empty and have less structure within them. We evaluate the genus measure both for individual voids as well as for the entire ensemble of voids predicted by CDM model. As a result we find that the topology of voids when taken together with the void spectrum is a very useful statistical indicator of the evolution of the structure of the universe on large scales.

Sahni, Varun↗

An Inverse Chance-constrained Approach to the Calibration of Robust Models

This paper proposes a strategy to calibrate computational models according to uncertain input-output data. To this end, uncertainty in the data is first quantified by creating adversarial data sets. Samples drawn from such sets are then mapped from the input-output space to the parameter space using an inverse mapping. This mapping minimizes the collective output spread of an ensemble of point predictions while satisfying a set of individual data-matching requirements. The distribution of the resulting parameter points, which often exhibits strong parameter dependencies, is then modeled using sliced-normals. The chance-constrained formulation used to learn this distribution enables the analyst to trade-off a greater likelihood for most of the data against a lower likelihood for some of the data thereby relaxing the conservatism of the calibrated model. This formulation not only neglects the worst-performing quantiles of each adversarial distribution but also eliminates the potentially serious effects that outliers might have on the resulting model. This calibration approach not only has a considerably lower computational cost than the standard forward approach but it also allows for the identification of suitable distribution classes, which in turn yield better calibrated models.

Calibration↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

Presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

This is the conference paper accompanying an oral presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?

Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties, including Monte Carlo dropout, deep ensemble, Bayesian neural networks, Gaussian Processes and conformal prediction. Lastly, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Seasonal Drought Prediction in East Africa: Can National Multi-Model Ensemble Forecasts Help?

The increasing food and water demands of East Africa's growing population are stressing the region's inconsistent water resources and rain-fed agriculture. As recently as in 2011 part of this region underwent one of the worst famine events in its history. Timely and skillful drought forecasts at seasonal scale for this region can inform better water and agro-pastoral management decisions, support optimal allocation of the region's water resources, and mitigate socio-economic losses incurred by droughts. However seasonal drought prediction in this region faces several challenges. Lack of skillful seasonal rainfall forecasts; the focus of this presentation, is one of those major challenges. In the past few decades, major strides have been taken towards improvement of seasonal scale dynamical climate forecasts. The National Centers for Environmental Prediction's (NCEP) National Multi-model Ensemble (NMME) is one such state-of-the-art dynamical climate forecast system. The NMME incorporates climate forecasts from 6+ fully coupled dynamical models resulting in 100+ ensemble member forecasts. Recent studies have indicated that in general NMME offers improvement over forecasts from any single model. However thus far the skill of NMME for forecasting rainfall in a vulnerable region like the East Africa has been unexplored. In this presentation we report findings of a comprehensive analysis that examines the strength and weakness of NMME in forecasting rainfall at seasonal scale in East Africa for all three of the prominent seasons for the region. (i.e. March-April-May, July-August-September and October-November- December). Simultaneously we also describe hybrid approaches; that combine statistical approaches with NMME forecasts; to improve rainfall forecast skill in the region when raw NMME forecasts lack in skill.

Shukla, Shraddhanand↗

Seasonal Drought Prediction in East Africa: Can National Multi-Model Ensemble Forecasts Help?

The increasing food and water demands of East Africa's growing population are stressing the region's inconsistent water resources and rain-fed agriculture. As recently as in 2011 part of this region underwent one of the worst famine events in its history. Timely and skillful drought forecasts at seasonal scale for this region can inform better water and agro-pastoral management decisions, support optimal allocation of the region's water resources, and mitigate socio-economic losses incurred by droughts. However seasonal drought prediction in this region faces several challenges. Lack of skillful seasonal rainfall forecasts; the focus of this presentation, is one of those major challenges. In the past few decades, major strides have been taken towards improvement of seasonal scale dynamical climate forecasts. The National Centers for Environmental Prediction's (NCEP) National Multi-model Ensemble (NMME) is one such state-of-the-art dynamical climate forecast system. The NMME incorporates climate forecasts from 6+ fully coupled dynamical models resulting in 100+ ensemble member forecasts. Recent studies have indicated that in general NMME offers improvement over forecasts from any single model. However thus far the skill of NMME for forecasting rainfall in a vulnerable region like the East Africa has been unexplored. In this presentation we report findings of a comprehensive analysis that examines the strength and weakness of NMME in forecasting rainfall at seasonal scale in East Africa for all three of the prominent seasons for the region. (i.e. March-April-May, July-August-September and October-November- December). Simultaneously we also describe hybrid approaches; that combine statistical approaches with NMME forecasts; to improve rainfall forecast skill in the region when raw NMME forecasts lack in skill.

Shukla, Shraddhanand↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)↗

EnZymClass: Substrate specificity prediction tool of plant acyl-ACP thioesterases based on ensemble learning

Characterizing the functional properties of plant acyl-ACP thioesterases (TEs), a key enzyme class used in the production of renewable oleochemicals in microbial hosts, experimentally, can be an expensive and time consuming process since it requires manual screening of thousands of candidates in a database. Using amino acid sequence to computationally predict an enzyme’s function might accelerate this process; however obtaining the necessary amount of information on previously characterized enzymes and their respective sequences required by standard Machine Learning (ML) based approaches to accurately infer sequence-function relationships can be prohibitive, especially with a low-throughput testing cycle. Experimental noise, unbalanced dataset where high sequence similarity does not always imply identical functional properties will further prevent robust prediction performance. Herein we present a ML method, Ensemble method for enZyme Classification (EnZymClass), that is specifically designed to address these issues. We used EnZymClass to classify TEs into short, long and mixed free fatty acid substrate specificity categories. While general guidelines for inferring substrate specificity have been proposed before, prediction of chain-length preference from primary sequence has remained elusive for plant acyl-ACP TEs. By applying EnZymClass to a subset of TEs in the ThYme database, we identified two medium chain TEs, ClFatB3 and CwFatB2, with previously uncharacterized activity in E. coli fatty acid production hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Interactive Vegetation Phenology, Soil Moisture, and Monthly Temperature Forecasts

The time scales that characterize the variations of vegetation phenology are generally much longer than those that characterize atmospheric processes. The explicit modeling of phenological processes in an atmospheric forecast system thus has the potential to provide skill to subseasonal or seasonal forecasts. We examine this possibility here using a forecast system fitted with a dynamic vegetation phenology model. We perform three experiments, each consisting of 128 independent warm-season monthly forecasts: 1) an experiment in which both soil moisture states and carbon states (e.g., those determining leaf area index) are initialized realistically, 2) an experiment in which the carbon states are prescribed to climatology throughout the forecasts, and 3) an experiment in which both the carbon and soil moisture states are prescribed to climatology throughout the forecasts. Evaluating the monthly forecasts of air temperature in each ensemble against observations, as well as quantifying the inherent predictability of temperature within each ensemble, shows that dynamic phenology can indeed contribute positively to subseasonal forecasts, though only to a small extent, with an impact dwarfed by that of soil moisture.

Seasonal Forecasting↗

Evaluation of the Four-Dimensional Ensemble-Variational Hybrid Data Assimilation with Self-Consistent Regional Background Error Covariance for Improved Hurricane Intensity Forecasts

The feasibility of a hurricane initialization framework based on the Gridpoint Statistical Interpolation (GSI)-based four-dimensional ensemble-variational (GSI-4DEnVar) hybrid data assimilation system for the Hurricane Weather Research and Forecasting model (HWRF) model is evaluated in this study. The system considers the temporal evolution of error covariances via the use of four-dimensional ensemble perturbations that are provided by high-resolution, self-consistent HWRF ensemble forecasts. It is different from the configuration of the GSI-based three-dimensional ensemble-variational (GSI-3DEnVar) hybrid data assimilation system, similar to that used in the operational HWRF, which employs background error covariances provided by coarser-resolution global ensembles from the National Centers for Environmental Prediction (NCEP) Global Forecast System (GFS) ensemble Kalman filtering data assimilation system. In addition, our proposed initialization framework discards the empirical intensity correction in the vortex initialization package that is employed by the GSI-3DEnVar initialization framework in operational HWRF. Data assimilation and numerical simulation experiments for Hurricanes Joaquin (2015), Patricia (2015), and Matthew (2016) are conducted during their intensity changes. The impacts of two initialization frameworks on the HWRF analyses and forecasts are compared. It is found that GSI-4DEnVar leads to a reduction in track, minimum sea level pressure (MSLP), and maximum surface wind (MSW) forecast errors in all of the HWRF simulations, compared with the GSI-3DEnVar initialization framework. With assimilating high-resolution observations within the hurricane inner-core region, GSI-4DEnVar can produce the initial hurricane intensity reasonably well without the empirical vortex intensity correction. Further diagnoses with Hurricane Joaquin indicate that GSI-4DEnVar can significantly alleviate the imbalances in the initial conditions and enhance the performance of the data assimilation and subsequent hurricane intensity and precipitation forecasts

54 ENVIRONMENTAL SCIENCES↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗