Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian parameter estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

How data science methods can improve the quality and efficiency of ICF and HEDP research

Data Science methods (many that are Bayesian based) are widely used in the physical sciences to estimate model parameters from experimental data, synthesize heterogeneous data, calibrate models, design experiments, and determine statistical significance of data. These methods provide a wealth of advantages over traditional analysis techniques because: 1) uncertainties are rigorously defined and propagated naturally through complex systems including covariance, 2) prior information is captured within the analysis framework (including rad-MHD and rad-hydro simulations), 3) competing models can be selected and/or ruled out using quantitative criteria, and 4) complex, heterogeneous data can be incorporated simultaneously. While these methods have been widely adopted as the gold standard in fields such as particle physics, astronomy, and biology, they have been slow to catch on in Inertial Confinement Fusion (ICF) and High Energy Density Physics (HEDP) research. Recently, several teams at LLNL, SNL, LANL, and the LLE have been exploring the use of these tools in their research and have found success. Here we propose that a concerted effort to consolidate these independent research efforts by developing and deploying common tools for use across the complex can revolutionize the way we approach data analysis, assimilation of theory and experiment, and decision making. The Bayesian formalism provides a means to accomplish this, but we are lacking certain infrastructure to make it happen on a large scale. Furthermore, once adopted, these techniques can be used to develop standards by which discoveries can be judged, similar to the so-called 5σ rule in high energy particle physics. Such standards may be used in the future to address the issue of unknown reproducibility in ICF and HED experiments caused by low shot rate and high cost per experiment. Our goals as a group are to advance the state of the art in HED measurement science by enabling: 1) better inferences from data with well-defined uncertainties, 2) better use of the data we have and continue to collect, 3) intelligent synthesis of data, 4) evaluation of the statistical significance of our data, and 5) informed decision making regarding the design of new experiments and instruments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

On the quantification and efficient propagation of imprecise probabilities with copula dependence

This paper addresses the problem of quantification and propagation of uncertainties associated with dependence modeling when data for characterizing probability models are limited. Practically, the system inputs are often assumed to be mutually independent or correlated by a multivariate Gaussian distribution. However, this subjective assumption may introduce bias in the response estimate if the real dependence structure deviates from this assumption. In this work, we overcome this limitation by introducing a flexible copula dependence model to capture complex dependencies. Here, a hierarchical Bayesian multimodel approach is proposed to quantify uncertainty in dependence model-form and model parameters that result from small data sets. This approach begins by identifying, through Bayesian multimodel inference, a set of candidate marginal models and their corresponding model probabilities, and then estimating the uncertainty in the copula-based dependence structure, which is conditional on the marginals and their parameters. The overall uncertainties integrating marginals and copulas are probabilistically represented by an ensemble of multivariate candidate densities. A novel importance sampling reweighting approach is proposed to efficiently propagate the overall uncertainties through a computational model. Through an example studying the influence of constituent properties on the out-of-plane properties of transversely isotropic E-glass fiber composites, we show that the composite property with copula-based dependence model converges to the true estimate as data set size increases, while an independence or arbitrary Gaussian correlation assumption leads to a biased estimate.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Preliminary Results on Bayesian Inverse UQ for OECD/NEA WPNCS Subgroup 14 Benchmark Exercise for Error Recovery and Experimental Coverage

The Organization for Economic Cooperation and Development (OECD) Working Party on Nucelar Criticality Safety (WPNCS) has proposed a benchmark exercise representative of neutronic behavior in criticality experiments. Here, the goal is to develop confidence in data assimilation techniques used to adjust nuclear data. Participants are given synthetic experimental models with associated measured data and asked to estimate the model parameters given the model and measurements as well as provide predictions for separate application models. In this work, we performed data assimilation using Bayesian inverse Uncertainty Quantification (UQ) with machine learning surrogate models to produce posterior parameter distributions for the requested parameters and posterior predictive distributions for the requested responses. Several experimental models are shown to insufficiently inform the posterior parameter distributions for the applications involved. However, given sufficient experimental data, posterior parameter estimates yielded reduced uncertainty in the response predictions of interest while covering the experimental data.

Bayesian Inference↗

Imprecise global sensitivity analysis using bayesian multimodel inference and importance sampling

Global Sensitivity Analysis (GSA) aims to understand the relative importance of uncertain input variables to model response. Conventional GSA involves calculating sensitivity (Sobol’) indices for a model with known model parameter distributions. However, model parameters are affected by aleatory and epistemic uncertainty, with the latter often caused by lack of data. In this paper, we propose a new framework to quantify uncertainty in probability model-form and model parameters resulting from small datasets and integrate these uncertainties into Sobol’ index estimates. First, the process establishes, through Bayesian multimodel inference, a set of candidate probability models and their associated probabilities. Imprecise Sobol’ indices are calculated from these probability models using an importance sampling reweighting approach. This results in probabilistic Sobol’ indices, whose distribution characterizes uncertainty in the sensitivity resulting from small dataset size. The imprecise Sobol’ indices thus provide a measure of confidence in the sensitivity estimate and, moreover, can be used to inform data collection efforts targeted to minimize the impact of uncertainties. Through an example studying the parameters of a Timoshenko beam, we show that these probabilistic Sobol’ indices converge to the true/deterministic Sobol’ indices as the dataset size increases and hence, distribution-form uncertainty reduces. The approach is then applied to assess the sensitivity of the out-of-plane properties of an E-glass fiber composite material to its constituent properties. This second example illustrates the approach for an important class of materials with wide-ranging applications when data may be lacking for some input parameters.

42 ENGINEERING↗

Hierarchical Bayesian Inverse Problems: A High-Dimensional Statistics Viewpoint

This paper analyzes hierarchical Bayesian inverse problems using techniques from highdimensional statistics. Furthermore, our analysis leverages a property of hierarchical Bayesian regularizers that we call approximate decomposability to obtain non-asymptotic bounds on the reconstruction error attained by maximum a posteriori estimators. The new theory explains how hierarchical Bayesian models that exploit sparsity, group sparsity, and sparse representations of the unknown parameter can achieve accurate reconstructions in high-dimensional settings.

MAP estimation↗

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)↗

Development and application of marginal likelihood optimization for integral parameter adjustment

When adjusting nuclear data with integral experiments, care must be taken that spurious adjustments are not made by assimilating poorly characterized integral parameters. If there are unaccounted for biases or poorly estimated uncertainties in the calculated and experimental values for an integral parameter, the Bayesian data assimilation may adjust the nuclear data in a manner that does not reflect the physics of the integral parameter. To identify and lessen the impact of these inconsistent integral parameters, in this study we present a Marginal Likelihood Optimization algorithm. In a data-driven way, the marginalized likelihood is used to modulate hyperparameter terms that decrease the influence of inconsistent integral parameters on the adjustment. The advantage of this approach over other methods in the literature is that it incorporates correlation information and does not remove an integral parameter from the adjustment. Herein, we present and motivate the algorithm, and apply it to an integral data assimilation case study.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

NEXTorch: A Design and Bayesian Optimization Toolkit for Chemical Sciences and Engineering

Automation and optimization of chemical systems require well-informed decisions on what experiments to run to reduce time, materials, and/or computations. Data-driven active learning algorithms have emerged as valuable tools to solve such tasks. Bayesian optimization, a sequential global optimization approach, is a popular active-learning framework. Past studies have demonstrated its efficiency in solving chemistry and engineering problems. Here we introduce NEXTorch, a library in Python/PyTorch, to facilitate laboratory or computational design using Bayesian optimization. NEXTorch offers fast predictive modeling, flexible optimization loops, visualization capabilities, easy interfacing with legacy software, and multiple types of parameters and data type conversions. It provides GPU acceleration, parallelization, and state-of-the-art Bayesian optimization algorithms and supports both automated an d human-in-the-loop optimization. The comprehensive online documentation introduces Bayesian optimization theory and several examples from catalyst synthesis, reaction condition optimization, parameter estimation, and reactor geometry optimization. NEXTorch is open-source and available on GitHub

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Probabilistic estimation of depth-resolved profiles of soil thermal diffusivity from temperature time series

Abstract. Improving the quantification of soil thermal and physical properties is key to achieving a better understanding and prediction of soil hydro-biogeochemical processes and their responses to changes in atmospheric forcing. Obtaining such information at numerous locations and/or over time with conventional soil sampling is challenging. The increasing availability of low-cost, vertically resolved temperature sensor arrays offers promise for improving the estimation of soil thermal properties from temperature time series, and the possible indirect estimation of physical properties. Still, the reliability and limitations of such an approach need to be assessed. In the present study, we develop a parameter estimation approach based on a combination of thermal modeling, sliding time windows, Bayesian inference, and Markov chain Monte Carlo simulation to estimate thermal diffusivity and its uncertainty over time, at numerous locations and at an unprecedented vertical spatial resolution (i.e., down to 5 to 10 cm vertical resolution) from soil temperature time series. We provide the necessary framework to assess under which environmental conditions (soil temperature gradient, fluctuations, and trend), temperature sensor characteristics (bias and level of noise), and deployment geometries (sensor number and position) soil thermal diffusivity can be reliably inferred. We validate the method with synthetic experiments and field studies. The synthetic experiments show that in the presence of median diurnal fluctuations ≥ 1.5 ∘C at 5 cm below the ground surface, temperature gradients > 2 ∘C m−1, and a sliding time window of at least 4 d the proposed method provides reliable depth-resolved thermal diffusivity estimates with percentage errors ≤ 10 % and posterior relative standard deviations ≤ 5 % up to 1 m depth. Reliable thermal diffusivity under such environmental conditions also requires temperature sensors to be spaced precisely (with accuracy to a few millimeters), with a level of noise ≤ 0.02 ∘C, and with a bias defined by a standard deviation ≤ 0.01 ∘C. Finally, the application of the developed approach to field data indicates significant repeatability in results and similarity with independent measurements, as well as promise in using a sliding time window to estimate temporal changes in soil thermal diffusivity, as needed to potentially capture changes in bulk density or water content.

54 ENVIRONMENTAL SCIENCES↗

Determining reference standard strength for neutron-irradiated reduced activation ferritic/martensitic steel F82H by Bayesian method

The deterministic approach widely adopted in the design of structural components relies on systematically defined design limits using empirically determined safety factors. However, this approach is not always appropriate because structures are subjected to a variety of loads in the practical environment, which may result in excessively conservative design limits. In recent years, a more rigorous probabilistic approach that incorporates material strength distributions has become an important solution. In the probabilistic approach, the probability density functions of material strength properties underpin the design criteria. Here, the objective of this study is to identify the density distribution functions that best describe tensile properties of irradiated F82H to define a reference strength for DEMO design. Due to the limited number of existing data, this study specifically employs a Bayesian prediction method based on Monte Carlo simulations to determine a material reference value with statistical reliability and to investigate its effectiveness. For example, the dependence of tensile properties of 300 °C irradiated materials on irradiation damage and the range predicted by 95% Bayesian estimation was evaluated. As a statistical model for the dose dependence of statistical parameters, the normal distribution exhibited a better fit for 0.2% proof strength and tensile strength, whereas the distribution of total elongation data gave comparable reference values for both the normal and Weibull distribution models. Both models gave comparable criteria for the distribution of total elongation data. The Weibull model also gave better results for uniform elongation. The function best describing the model was a logarithmic law for both 0.2% proof strength and tensile strength, while a power law for both total and uniform elongation, which allowed for more comprehensive data prediction of irradiation data with statistical accuracy for DEMO reactor design.

36 MATERIALS SCIENCE↗

PRIME - A Software Toolkit for the Characterization of Partially Observed Epidemics in a Bayesian Framework

PRIME is a modeling framework designed for the “real-time’” characterization and forecasting of partially observed epidemics. Characterization is the estimation of infection spread parameters using daily counts of symptomatic patients. The method is designed to help guide medical resource allocation in the early epoch of the outbreak. The estimation problem is posed as one of Bayesian inference and solved using a Markov Chain Monte Carlo technique. The framework can accommodate multiple epidemic waves and can help identify different disease dynamics at the regional, state, and country levels. We include examples using publicly available COVID-19 data.

97 MATHEMATICS AND COMPUTING↗

Development of a framework for sequential Bayesian design of experiments: Application to a pilot-scale solvent-based CO 2 capture process

In this paper, a methodology is developed for sequential design of experiments (SDoE) for process systems and applied to a solvent-based CO 2 capture system. In this approach, the prior knowledge of the system is used to prioritize process data collection at specific operating conditions. These data are then incorporated into a Bayesian inference methodology for updating a stochastic model by refining estimations of its underlying parameters, and the updated model is then used to generate the next set of test runs. Thus, the new knowledge obtained from the data is used to guide subsequent iterations of the experimental runs, ensuring that the overall data collection is maximally informative given that most experimental campaigns, especially at pilot or higher-scale plants, are costly, time-consuming, and resource-limited. The test run objective for this work was to minimize the maximum model prediction uncertainty for key output variables, but the methodology is generic and can be readily applied to other test run objectives. This methodology is applied to an aqueous monoethanolamine (MEA) pilot plant campaign at the National Carbon Capture Center (NCCC) in Wilsonville, Alabama, USA. The SDoE framework was utilized for two iterations, while collecting 18 sets of data representing different process conditions, and this resulted in an overall average reduction in uncertainty of approximately 50% in the prediction of CO 2 capture percentage. Moreover, 11 additional data sets were obtained with variation of absorber packing height for further model validation. This work shows the capability of the SDoE framework to maximize learning given limited resources, allowing for the reduction of model uncertainty, which is of great importance for many applications including reduction of technical risk associated with scale-up and economic analysis.

20 FOSSIL-FUELED POWER PLANTS↗

Characterization of Partially Observed Epidemics - Application to COVID-19

This report documents a statistical method for the "real-time" characterization of partially observed epidemics. Observations consist of daily counts of symptomatic patients, diagnosed with the disease. Characterization, in this context, refers to estimation of epidemiological parameters that can be used to provide short-term forecasts of the ongoing epidemic, as well as to provide gross information for the time-dependent infection rate. The characterization problem is formulated as a Bayesian inverse problem, and is predicated on a model for the distribution of the incubation period. The model parameters are estimated as distributions using a Markov Chain Monte Carlo (MCMC) method, thus quantifying the uncertainty in the estimates. The method is applied to the COVID-19 pandemic of 2020, using data at the country, provincial (e.g., states) and regional (e.g. county) levels. The epidemiological model includes a stochastic component due to uncertainties in the incubation period. This model-form uncertainty is accommodated by a pseudo-marginal Metropolis-Hastings MCMC sampler, which produces posterior distributions that reflect this uncertainty. We approximate the discrepancy between the data and the epidemiological model using Gaussian and negative binomial error models; the latter was motivated by the over-dispersed count data. For small daily counts we find the performance of the calibrated models to be similar for the two error models. For large daily counts the negative-binomial approximation is numerically unstable unlike the Gaussian error model. Application of the model at the country level (for the United States, Germany, Italy, etc.) generally provided accurate forecasts, as the data consisted of large counts which suppressed the day-to-day variations in the observations. Further, the bulk of the data is sourced over the duration before the relaxation of the curbs on population mixing, and is not confounded by any discernible country-wide second wave of infections. At the state-level, where reporting was poor or which evinced few infections (e.g., New Mexico), the variance in the data posed some, though not insurmountable, difficulties, and forecasts were able to capture the data with large uncertainty bounds. The method was found to be sufficiently sensitive to discern the flattening of the infection and epidemic curve due to shelter-in-place orders after around 90% quantile for the incubation distribution (about 10 days for COVID-19). The proposed model was also used at a regional level to compare the forecasts for the central and north-west regions of New Mexico. Modeling the data for these regions illustrated different disease spread dynamics captured by the model. While in the central region the daily counts peaked in the late April, in the north-west region the ramp-up continued for approximately three more weeks.

59 BASIC BIOLOGICAL SCIENCES↗

Characterization of partially observed epidemics through Bayesian inference: application to COVID-19

We demonstrate a Bayesian method for the "real-time'" characterization and forecasting of partially observed COVID-19 epidemic. Characterization is the estimation of infection spread parameters using daily counts of symptomatic patients.The method is designed to help guide medical resource allocation in the early epoch of the outbreak. The estimation problem is posed as one of Bayesian inference and solved using a Markov chain Monte Carlo technique. The data used in this study was sourced before the arrival of the second wave of infection in July 2020. The proposed modeling approach, when applied at the country level, generally provides accurate forecasts at the regional, state and country level. The epidemiological model detected the flattening of the curve in California, after public health measures were instituted.The method also detected different disease dynamics when applied to specific region of New Mexico

60 APPLIED LIFE SCIENCES↗

Machine learning based approach to predict ductile damage model parameters for polycrystalline metals

Damage models for ductile materials typically need to be parameterized, often with the appropriate parameters changing for a given material depending on the loading conditions. This can make parameterizing these models computationally expensive, since an inverse problem must be solved for each loading condition. Using standard inverse modeling techniques typically requires hundreds or thousands of high-fidelity computer simulations to estimate the optimal parameters. Additionally, the time of a human expert is required to set up the inverse model. Machine learning has recently emerged as an alternative approach to inverse modeling in these settings, where the machine learning model is trained in an offline manner and new parameters can be quickly generated on the fly, after training is complete. Here, this work utilizes such a workflow to enable the rapid parameterization of a ductile damage model called TEPLA with a machine learning inverse model. The machine learning model can efficiently estimate the model parameters much faster, as compared to previously employed methods, such as Bayesian calibration. The results demonstrate good accuracy on a synthetic test dataset and is validated against experimental data.

36 MATERIALS SCIENCE↗

Fast, differentiable, and extensible big bang nucleosynthesis package

Here, we introduce light isotope nucleosynthesis with JAX (LINX), a new differentiable public big bang nucleosynthesis code designed for fast parameter estimation. By leveraging JAX, LINX achieves both speed and differentiability, enabling the use of Bayesian inference, including gradient-based methods. We discuss the formalism used in LINX for rapid primordial elemental abundance predictions and give examples of how LINX can be used. When combined with differentiable cosmic microwave background power spectrum emulators, LINX can be used for joint cosmic microwave background and big bang nucleosynthesis analyses without requiring extensive computational resources, including on personal hardware.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PRIME

SAND2021-0565 O PRIME is a modeling framework designed for the real-time characterization and forecasting of partially observed epidemics. The method is designed to help guide medical resource allocation in the early epoch of the outbreak. Characterization is the estimation of infection spread parameters using daily counts of symptomatic patients. The estimation problem is posed as one of Bayesian inference and solved using a Markov Chain Monte Carlo technique. The framework can accommodate multiple epidemic waves and can help identify different disease dynamics at the regional, state, and country levels. Examples are provided using publicly available COVID-19 data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

Multi-Modal Bayesian Neural Network Surrogates with Conjugate Last-Layer Estimation

As data collection and simulation capabilities advance, multi-modal learning, the task of learning from multiple modalities and sources of data, is becoming an increasingly important area of research. Surrogate models that learn from data of multiple auxiliary modalities to support the modeling of a highly expensive quantity of interest have the potential to aid outer loop applications such as optimization, inverse problems, or sensitivity analyses when multi-modal data are available. We develop two multi-modal Bayesian neural network surrogate models and leverage conditionally conjugate distributions in the last layer to estimate model parameters using stochastic variational inference (SVI). We provide a method to perform this conjugate SVI estimation in the presence of partially missing observations. Here, we demonstrate improved prediction accuracy and uncertainty quantification compared to unimodal surrogate models for both scalar and time series data.

97 MATHEMATICS AND COMPUTING↗