AI Surrogate Model for Distributed Computing Workloads
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A primary advantage to using reduced complexity climate models (RCMs) has been their ability to quickly conduct probabilistic climate projections, a key component of uncertainty quantification in many impact studies and multisector systems. Providing frameworks for such analyses has been a target of several RCMs used in studies of the future co-evolution of the human and Earth systems. In this paper, we present Matilda, an open-science R software package that facilitates probabilistic climate projection analysis, implemented here using the Hector simple climate model in a seamless and easily applied framework. The primary goal of Matilda is to provide the user with a turn-key method to build parameter sets from literature-based prior distributions, run Hector iteratively to produce perturbed parameter ensembles (PPEs), weight ensembles for realism against observed historical climate data, and compute probabilistic projections for different climate variables. This workflow gives the user the ability to explore viable parameter space and propagate uncertainty to model ensembles with just a few lines of code. The package provides significant freedom to select different scoring criteria and algorithms to weight ensemble members, as well as the flexibility to implement custom criteria. Additionally, the architecture of the package simplifies the process of building and analyzing PPEs without requiring significant programming expertise, to accommodate diverse use cases. We present a case study that provides illustrative results of a probabilistic analysis of mean global surface temperature as an example of the software application.
This report documents work sponsored by the U.S. Nuclear Regulatory Commission (NRC) at the Oak Ridge National Laboratory (ORNL) as part of the RES project, “Application of Point Precipitation Frequency Estimates to Watersheds.” This project was implemented as part of the Probabilistic Flood Hazard Assessment (PFHA) Research Program. The objective of the PFHA Research Program is to develop tools and guidance on the use of PFHA methods to risk-inform NRC’s licensing of new facilities as well as licensing and oversight of currently operating facilities as they relate to flooding hazards. Many nuclear power plants (NPPs) are located on or near rivers so riverine flooding hazards need to be considered in their design and operation. Probabilistic riverine flood models are important tools for realistic assessment of flooding risks. However, these models require areal estimates of the depth, duration, and frequency of rainfall distributed over the watershed, which are not often available. Point precipitation frequency estimates are more widely available. For example, the National Oceanic and Atmospheric Administration (NOAA) has published NOAA Atlas 14, which provides point precipitation frequency estimates for 5-minute through 60-day durations at average recurrence intervals of 1-year through 1,000-year. The research documented in this report addresses areal reduction factors (ARFs), which can be used to convert the widely available point precipitation frequency estimates, to estimates of areal precipitation frequency over a watershed. The most widely used ARF source is Technical Paper 29 (TP-29) published by the then U.S. Weather Bureau in 1958. However, both the methods and the underlying precipitation data used to produce TP-29 are seriously out of date. For example, due to the small gauge network available at the time of TP-29’s compilation, ARF estimates developed are only for watersheds smaller than about 400 square miles. Due to the relatively short record lengths of precipitation data available, frequency considerations could not be accurately determined. Other factors such as regional climate and seasonality were not addressed. Several newer methods have been published since TP-29 was developed and both the type and quantity of precipitation data have increased significantly, along with computational resources and analytical tools such as geographic information systems. This report reviewed and assessed the available precipitation products and methods for conducting ARF analysis. The work applied up-to-date precipitation data products and analysis methods with a novel watershed-based approach to investigate how ARF estimates vary across different methods, data sources, geographical locations, return periods, and seasons. The overall findings reported here regarding basic ARF trends are in line with other recent studies showing that ARFs decrease with increasing area, increase with increasing duration, and decrease with increasing return period. This study found significant differences among the available ARF methods. This work also found a strong geographical variability across different US hydrologic regions, suggesting that the ARF are specific to regional climate patterns and geographical characteristics and should not be applied arbitrarily to other locations. The results also reveal the importance of data record length, especially for high return level ARFs. The work reported in NUREG/CR-7271 will assist NRC staff in assessing different classes of ARF methods in conjunction with available rainfall data sets. It will also support the development of guidance for application of point precipitation data in PFHAs. It should be noted that the ARF values presented in this report for any location or region were developed for the purposes of comparing methods and investigating the factors that influence ARFs. They should not be considered official and should not be used in leu of a site-specific analysis.
The stable numerical integration of shocks in compressible flow simulations relies on the reduction or elimination of Gibbs phenomena (unstable, spurious oscillations). A popular method to virtually eliminate Gibbs oscillations caused by numerical discretization in under-resolved simulations is to use a flux limiter. A wide range of flux limiters have been studied in the literature, with recent interest in their optimization via machine learning methods trained on high-resolution datasets. The common use of flux limiters in numerical codes as plug-and-play blackbox components makes them key targets for design improvement. Even for deterministic dynamical models, numerical uncertainty is introduced via coarse-graining required by insufficient computational power to solve all scales of motion. Conventional flux limiters are deterministic and lack the capacity to address uncertainties, both aleatoric (inherent randomness) and epistemic (modeling uncertainty due to limited knowledge), which arise in coarse-grained numerical simulations. Here, we introduce a conceptually distinct type of flux limiter that is designed to handle the effects of randomness in the model and uncertainty in model parameters. Unlike traditional single-function flux limiters, these new probabilistic flux limiters incorporate multiple flux limiting functions, each applied with a learned probability drawn from high-resolution data to mitigate the effects of uncertainty in numerical simulations. This approach departs from traditional single-function limiters by explicitly modeling and incorporating uncertainty into the shock capturing process. Using the example of Burgers' equation as a testbed, we show that a machine learned, probabilistic flux limiter may be used in a shock capturing code to more accurately capture shock profiles. In particular, we show that our probabilistic flux limiter outperforms standard limiters and can be successively improved upon (up to a point) by expanding the set of probabilistically chosen flux limiting functions.
This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.
Net load imbalances from day ahead forecasts can lead to significant grid operations costs and are expected to increase as variable renewable energy adoption grows. We propose a new wholesale market product to manage the risk of net load imbalances called Flexibility Options. This product relies on probabilistic forecasts to estimate flexibility demand and would be co-optimized in the day-ahead market. We also propose stochastic methods that enable DER and flexible load aggregators to participate in flexibility markets while considering the uncertainty in weather and occupant behavior.
Satellite data provides essential insights into the spatiotemporal distribution of CO 2 concentrations. However, many atmospheric inverse models fail to adequately incorporate the spatial and temporal correlations inherent in satellite observations and often lack rigorous methods for estimating parameters like spatial length scales. We introduce an inference model that processes the spatiotemporal covariance in satellite data and estimates hyperparameters such as covariance length scales. Our approach uses the Gaussian process (GP) machine learning (ML) and modern probabilistic programming languages (PPLs) to perform atmospheric inversions of emissions from satellite data. We develop a GP ML inversion system based on modern PPLs and the GEOS-Chem chemical transport model, simulating atmospheric CO 2 concentrations corresponding to the Orbiting Carbon Observatory-2/3 (OCO-2/3) data for July 2020. In our supervised learning framework, we treat the GEOS-Chem simulated data set as the target, with predictors derived by scaling the target with sector-specific factors hidden from the GP machine. Our results show that the GP model, combined with GPU-enabled PPLs, effectively retrieves true emission scaling factors and infers noise levels concealed within the data. This suggests that our method could be applied over larger areas with more complex covariance structures, enabling comprehensive analysis of the spatiotemporal patterns observed in OCO-2/3 and similar satellite data sets.
This paper proposes an analytic neural network Gaussian process (NNGP)-based chance-constrained real-time voltage regulation method for active distribution systems with photovoltaics (PVs), batteries, and electric vehicles (EVs). NNGP can utilize historical measurement data to achieve real-time probabilistic node voltage estimation through Bayesian inference. Then, NNGP is fully analytically embedded into the optimal power flow model to perform voltage regulation and adapt to various topological changes. The uncertainties of voltage estimations are easily considered via the chance constraint, and it has been shown that the adoption of this chance constraint can significantly improve the reliability of voltage regulation under various scenarios. The comparison results with other methods, carried out on a real 759-node distribution system located in western Colorado, U.S., show that the proposed method can achieve accurate voltage estimation across different topologies and reliably perform voltage regulation considering PVs, batteries, and EVs.
We introduce here an approach based on the Givens representation for posterior inference in statistical models with orthogonal matrix parameters, such as factor models and probabilistic principal component analysis (PPCA). We show how the Givens representation can be used to develop practical methods for transforming densities over the Stiefel manifold into densities over subsets of Euclidean space. We demonstrate how to deal with issues arising from the topology of the Stiefel manifold and how to inexpensively compute the change-of-measure terms. We introduce an auxiliary parameter approach that limits the impact of topological issues. We provide both analysis of our methods and numerical examples demonstrating the effectiveness of the approach. We also discuss how our Givens representation can be used to define general classes of distributions over the space of orthogonal matrices. We then give demonstrations on several examples showing how the Givens approach performs in practice in comparison with other methods.
In general, Monte Carlo methods simulate large numbers of random trials in order to observe numerical behavior of systems described by probabilistic behavior. In radiation transport, pseudo-random number generators are used to randomly sample individual particle lives. Information about the particles are tallied.
We present differentiable predictive control (DPC), a method for offline learning of constrained neural control policies for nonlinear dynamical systems with performance guarantees. We show that the sensitivities of the parametric optimal control problem can be used to obtain direct policy gradients. Specifically, we employ automatic differentiation (AD) to efficiently compute the sensitivities of the model predictive control (MPC) objective function and constraints penalties. To guarantee safety upon deployment, we derive probabilistic guarantees on closed-loop stability and constraint satisfaction based on indicator functions and Hoeffding’s inequality. We empirically demonstrate that the proposed method can learn neural control policies for various parametric optimal control tasks. In particular, we show that the proposed DPC method can stabilize systems with unstable dynamics, track time-varying references, and satisfy nonlinear state and input constraints. Our DPC method has practical time savings compared to alternative approaches for fast and memory-efficient controller design. Specifically, DPC does not depend on a supervisory controller as opposed to approximate MPC based on imitation learning. We demonstrate that, without losing performance, DPC is scalable with greatly reduced demands on memory and computation compared to implicit and explicit MPC while being more sample efficient than model-free reinforcement learning (RL) algorithms.
Uncertainty in predicting solar energy resources introduces major challenges in power system management and necessitates the development of reliable probabilistic solar forecasts. As the first part of the development of probabilistic forecasts based on the Weather Research and Forecasting model with solar extensions (WRF-Solar), this study presents a tangent linear approach to identify input variables responsible for the largest uncertainties in predicting surface solar irradiance and clouds. A tangent linear analysis is capable of efficiently investigating sensitivities of output variables with respect to various input variables of WRF-Solar because this approach avoids the computational burden of perturbing the initial conditions of individual input variables. We develop tangent linear models (TLMs) for six WRF-Solar physics packages that control the formation and dissipation of clouds and solar radiation, and we evaluate the validity of TLMs using a linearity test. The tangent linear sensitivity analysis is conducted under various scenarios based on satellite observations and model simulations to consider realistic input conditions. A simple method is used to quantify the impact of the uncertainty of input variables on the output variables from the TLMs. The results demonstrate that uncertainties in the output variables that are the focus of this study—including global horizontal irradiance, direct normal irradiance, cloud mixing ratio, cloud tendency, cloud fraction, and sensible and latent heat fluxes—are highly sensitive to uncertainties in 14 input variables. This study indicates that the tangent linear method can identify key variables of physics modules in WRF-Solar that can be stochastically perturbed to generate ensemble-based probabilistic forecasts.
Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.
In internal combustion engine research, cylinder pressure measurements provide valuable information about the underlying thermodynamic and combustion processes, and are typically collected in ensembles of several 100 traces. Although in some particular fields of combustion research all traces are analyzed, in most cases only one trace is studied because analyzing all the traces is impractical due to the large number of collected samples. Instead, an ensemble-averaged pressure trace is commonly calculated and used for analysis. However, this pressure trace is highly smoothed and dynamic information is lost during the averaging process. With the average trace, pressure rise rates are lower and pressure oscillations such as the ones resulting from combustion knock are lost. In this work, a statistical method was developed to determine the “most representative cycle,” which is the cycle from the ensemble that has the pressure trace most representative of the engine operating condition. Eleven characteristic parameters are computed from each pressure trace and probabilistic distributions are obtained for each of the parameters using all the traces in the ensemble. Finally, the most representative cycle is selected by means of a cost function minimization. The benefits of this method are illustrated using experimental data from four very different engine platforms, under four different combustion modes and over a range of operating conditions.
The current and next observation seasons will detect hundreds of gravitational waves (GWs) from compact binary systems coalescence at cosmological distances. When combined with independent electromagnetic measurements, the source redshift will be known, and we will be able to obtain precise measurements of the Hubble constant H_0 via the distance–redshift relation. However, most observed mergers are not expected to have electromagnetic counterparts, which prevents a direct redshift measurement. In this scenario, one possibility is to use the dark sirens method that statistically marginalizes over all the potential host galaxies within the GW location volume to provide a probabilistic source redshift. Here we presented H_0 measurements using two new dark sirens compared to previous analyses using DECam data: GW190924|$\_$|021846 and GW200202|$\_$|154313. The photometric redshifts of the possible host galaxies of these two events are acquired from the DECam Local Volume Exploration Survey (DELVE) carried out on the Blanco telescope at Cerro Tololo. The combination of the H_0 posterior from GW190924|$\_$|021846 and GW200202|$\_$|154313 together with the bright siren GW170817 leads to |$H_{0} = 68.84^{+15.51}_{-7.74}\, \rm {km\, s^{-1}\, Mpc^{-1}}$|. Including these two dark sirens improves the 68 per cent confidence interval (CI) by 7 per cent over GW170817 alone. This demonstrates that the addition of well-localized dark sirens in such analysis improves the precision of cosmological measurements. Using a sample containing 10 well-localized dark sirens observed during the third LIGO/Virgo observation run, without the inclusion of GW170817, we determine a measurement of |$H_{0} = 76.00^{+17.64}_{-13.45}\, \rm {km\, s^{-1}\, Mpc^{-1}}$|.
Critical temperature for localized corrosion can be a good design parameter because localized corrosion is not likely to occur below that temperature. The critical temperature depends on alloy composition, microstructure, and environment chemistry (including its redox potential). This paper reviews the literature on critical temperature for localized corrosion, expressed either as Critical Pitting Temperature (CPT) or Critical Crevice Temperature (CCT). A history of various testing methods is presented. Different approaches for modeling the temperature of transition to active pit growth are reviewed, including probabilistic aspects of critical temperature. A semi-empirical, electrolyte-based, model is described that can be useful in predicting CCT in service environments that differ from standard laboratory test environments. The model predictions are compared to experimental data for various alloys. The effect of solvent on CCT/CPT is described briefly and future avenues of research are recommended.
Moment tensors (MTs) have long been used in earthquake and explosion source analysis, and there has been a renewed interest in how they can inform us about the seismic source, particularly in the geophysical monitoring community due to its application in event identification and yield analysis. However, parameter uncertainties in seismic MT inversion are rarely available. The inverse procedure often does not quantify MT model errors such as event location, data noise and Earth model that are essential for estimating solution robustness. To address this need, we propose to adopt the Bayesian probabilistic framework to incorporate uncertainties in MT inversions. In this study, we present the theoretical background of a probabilistic Bayesian framework for MT inversion accounting for model and measurements errors and illustrate the implementation of the method using a synthetic example.
The evidence of climate change is increasingly well-documented and impacts should be incorporated in performance assessment studies. The current climate literature provides both observational evidence and climate model projections of climate trends and/or climate change in the late 20. and early 21. centuries for North America and the northeast United States. Probabilistic modeling is a core requirement for quantifying uncertainty and evaluating its impacts. Not evaluating future climate states in a performance assessment because of the existence of uncertainty is contradictory to good modeling practices - the most uncertain issues and parameters require the most attention in effective probabilistic modeling. Excluding climate change limits development of modeling information that could aid in effective decision making. In this work we develop methods to use the output from hydrologic models and analysis of historical aerial imagery to quantify and implement the impacts of climate change on hydrologic processes at a nuclear waste site in West Valley, New York. Specifically, we used the HELP (Hydraulic Performance of Landfill Performance) model to characterize key hydrologic processes under both current and future climate conditions to assess the impacts of changing climate on hydrology. A suite of previous climatic models were reviewed and synthesized to produce a cohesive representation of the current state of knowledge of the impact of climate change on important model inputs such as precipitation. Output from the HELP simulations was coupled to the GoldSim model that was used to develop the Probabilistic Performance Assessment (PPA) approach through the application of a novel 'nearest neighbor' technique. First, several thousand realizations were generated from the HELP model using a Latin Hypercube experimental design to ensure adequate coverage of the parameter space of explanatory variables used to drive HELP. For example, porosity is a physical parameter that is used as an input to both HELP and the GoldSim PA model. We then ran sensitivity analysis (SA) algorithms on the output of HELP for each of the responses of interest. For each predictor, each time we build an SA model we get a different value for the sensitivity index (SI). From the collection of all the SA models, the average was computed among all of the SIs to represent the predictor within the context of the nearest neighbor approach. That is, we conduct SA on each HELP outcome for each scenario. This gives us parameter sensitivity indices for the outcomes. We average the parameter sensitivity indices across the outcomes to get the average SI for a scenario. For each realization that is generated from the Goldsim PA model, Goldsim generates random values for physical/empirical parameters that HELP uses as well. For each vector of physical/empirical parameters that Goldsim generates, the vector from the 5,000 HELP runs that is most 'similar' to the Goldsim vector is computed using the nearest neighbor approach. In this context 'similar' means minimization of the SA-weighted sum of the absolute differences among the 5,000 values computed for this statistic, where each value corresponds to a different HELP realization. In order to account for the impacts of climate change, this process was repeated using the spatially downscaled future climate projections. For each of the key parameters of interest, it was assumed that a linear change depicted the relationship between the values for the present day and those for 2100. In this way, the climatically-driven changes in key parameters used to inform the GoldSim model are quantified and incorporated into the PA model output for the future. (authors)