A collaborative study on the precision of the Markov chain Monte Carlo algorithms used for DNA profile interpretation
Not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This thesis reports a constraint of the neutrino oscillation parameters $\Delta m^{2}_{32}$, $\sin^2 \theta_{23}$, and $\delta_{CP}$ using the NuMI Off-Axis $\nu$ Appearance (NOvA) experiment's Near Detector (ND) data and Far Detector (FD) fake data set simultaneously. This thesis also reports a constraint on NOvA's systematic uncertainty model solely with its Near Detector data. The Hamiltonian Monte Carlo algorithm is used to estimate Bayesian Credible Intervals for the oscillation and interaction parameters. The $1\sigma$ Credible Intervals for $\sin^2 \theta_{23}$ are $(0.44, 0.512)$ $\cup$ $(0.536, 0.56)$, for $\Delta m^{2}_{32}$ $(2.41 \times 10^{-3}$ eV$^2,\ 2.52 \times 10^{-3}$ eV$^2)$, and for $\delta_{CP}$ $(0.74\pi,\ 1.1\pi)$ $\cup$ $(1.38\pi,\ 1.58\pi)$. The statistical power of the ND data constrains NOvA's interaction parameters, while the FD fake data constrains the oscillation parameters. This is the first analysis within NOvA to constrain the ND and FD prediction sim ultaneously, and to investigate the neutrino interaction modeling in the context of constraining the oscillation parameters. To constrain the ND data requires a sophisticated understanding of the neutrino interaction modeling and its uncertainties. The interested reader is advised to focus on Chapters 4 and 6, which discuss the ND selection, uncertainties, and ND-only fits to data. The reader interested in oscillation parameter constraints will find this in Chapter 7.
The primary goals of this project are exploring hidden geothermal resources in the U.S.A. and designing profitable enhanced geothermal systems (EGS). Many processes and parameters control geothermal exploration and energy production from geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize subsurface geothermal conditions. Sparse and multi-scale characteristics of these datasets prohibit properly leveraging these datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) promise to resolve these issues. The tremendous challenges and risks of geothermal exploration and production bring the demand for novel ML methods and tools that can (1) analyze large field datasets, (2) assimilate model simulations (large inputs and outputs), (3) process sparse datasets, (4) perform transfer learning (between sites with different exploratory levels), (5) extract hidden geothermal signatures in the field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. To address these necessities, ML-based geothermal resources exploration and enhanced geothermal systems (EGS) design tools have been developed. The exploration tool is called GeoThermalCloud and EGS design tool is called GeoDT-ML. GeoThermalCloud (https://github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. Also, it enables the identification of critical measurements needed to identify geothermal resource signatures. Alternatively, GeoDT-ML (https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS) is an ML-based alternative to GeoDT (https://github.com/GeoDesignTool/GeoDT.git), a fast, simplified multi-physics solver to evaluate EGS project designs in uncertain geologic systems. GeoDT-ML leverages recent advances in deep learning and high-performance computing. It is a faster and simpler version of GeoDT. To make this project a success, we used capabilities of LANL, PNNL, Google, Stanford, and Julia Computing. We analyzed eight datasets of the U.S.A. using GeothermalCloud and demonstrated potential highly prospective geothermal resources and identified key factors defining highly prospective sites. The first data set includes 44 locations in southwest New Mexico and 18 geological, hydrogeological, geophysical, geothermal, geochemical attributes. We defined low- and medium-temperature hydrothermal systems and discovered a new highly prospective site. The second data set analyzed 18 shallow water chemistry attributes at 14,342 locations in the Great Basin. It demarcated modestly, moderately, and highly prospective sites including key attributes for each type of prospectivity. The third data set analyzed Utah FORGE data including satellite (InSAR), geophysical (gravity, seismic), geochemical, and geothermal attributes. Here, we performed prospectivity analysis to identify future drilling locations using geological, geochemical, and geophysical attributes. Maps of temperature at depth and heat flow are constructed based on the available data. Prospectivity maps were generated, and drilling locations were proposed for future geothermal field exploration. The fourth data set analyzed 21 attributes at 120 locations in Tularosa Basin, New Mexico; data comes from past play fairway analyses in this region. ML analyses identified geothermal signatures associated with modestly, moderately, and highly hydrothermal systems. We also defined dominant attributes and spatial distribution of the geothermal signatures. The fifth, sixth, seventh, and eighth datasets include Tohatchi Springs, New Mexico, Hawaii, Brady site, Nevada, and EGS Collab, respectively. Moreover, we coupled GeothermalCloud and magnetotellurics data to pinpoint drilling locations for developing geothermal projects in the Tularosa Basin, New Mexico. GeothermalCloud found potential prospective locations for geothermal resources near White Sands Missile Range and McGregor Range at Fort Bliss. Magnetotellurics data determined the potential depth (~1800m) of geothermal prospects at McGregor Range based on apparent resistivity structures/layers in the subsurface. The McGregor Range consists of three resistivity layers and two resistivity structures. Magnetotellurics data also helps identify that the western portion of the McGregor Range has thick and low-resistivity earth materials. The low resistivity to the west is most likely for a fault system. Assuming temperature is consistent with a geothermal reservoir, the west-central part of the McGregor Range has the highest geothermal potential because of the increase in porosity and associated permeability attributed to the interpreted fault system. Also, we devised a coupling strategy between a process model and GeothermalCloud to characterize hydrogeological conditions and geothermal conditions, respectively. The process model characterizes hydrogeological and geothermal conditions on highly prospective geothermal sites provided by GeothermalCloud. We developed a physics-informed neural network (PINN) version of the Burns equation that can be easily coupled with GeothermalCloud. Furthermore, we performed an optimal design decision maximizing the economic value of an EGS power plant. This study optimized the range of well spacing between injection and production wells maximizing net present value in dollars (NPV). For this task, we used the GeoDT to simulate the Utah FORGE EGS development cycle from the initial well design to the end of production. Next, we accomplished another crucial task, which is predicting permeability of geothermal reservoirs. Predicting permeability of geothermal reservoirs is a non-trivial task because of huge computational runtime of simulation and lack of measurements. To avoid these limitations, we used easy-to-measure chemical concentrations in the subsurface as measurement data and convolutional neural network based ML model of a high-fidelity model. Next, we predicted permeability using Markov chain Monte Carlo simulation. We found that Markov chain Monte Carlo simulation predicts permeability with a high certainty if the prediction zone in the simulation area has chemical concentration data. Finally, we analyzed the DOE funded INGENIOUS and GeoDAWN projects data. For discovering hidden geothermal systems in the Great Basin, the INGENIOUS project accumulated old data, collected new data, and released them in 2022. The dataset includes a total of 24 geological, geophysical, and geochemical attributes. Data resolution and scale significantly vary prohibiting an appropriate usage. To avoid such limitations, we brought all data in the same resolution and scale by applying the inverse distance weighting interpolation technique for predicting data in unsampled locations. Subsequently, we analyzed LiDAR data of the GeoDAWN project. We received data in tiles format. The DOE’s overarching goal is to use ML on LiDAR data for finding favorable geological structures (e.g., step up faults in Brady, Nevada). To serve the purpose, we need to label favorable geologic structures that correspond to LiDAR data. We wrote an algorithm to label the LiDAR data with the favorable geologic structures.
MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.
Not Available
Background: The nuclear shell model is a powerful framework for predicting nuclear structure observables, but relies on interaction matrix elements fit to experimental data as its inputs. Extending the shell model's applicability, particularly toward dripline nuclei, requires efficient fitting methods and credible uncertainty quantification. Traditional approaches face computational challenges and may underestimate uncertainties. Purpose: We develop and test a framework combining eigenvector continuation and Markov chain Monte Carlo to efficiently fit shell model interaction matrix elements and quantify their uncertainties. Methods: Eigenvector continuation is used to emulate shell model calculations, reducing computational costs. The emulator enables Markov chain Monte Carlo sampling to optimize interaction matrix elements and rigorously assess parametric uncertainties. Here, the framework is benchmarked using the USDB interaction in the 𝑠𝑑 shell. Results: The emulator reproduces the USDB interaction with negligible error, validating its use in shell model fitting applications. However, we find that to obtain credible predictive intervals, the model defect of the shell model itself, rather than experimental or emulator error, must be taken into account in order to obtain credible uncertainties. Conclusions: The proposed framework provides an efficient and rigorous approach for fitting shell model interactions and quantifying uncertainties. Further, the normality assumption used in the past appears sufficient to describe the distribution of interaction matrix elements. However, it is crucial to account for model correlations to avoid underestimating uncertainties.
In recent work, we developed a Markov Chain Monte Carlo (MCMC) procedure to predict the ground state masses capable of forming the observed Solar r-process rare-earth abundance peak. By applying this method to nucleosynthesis calculations which make use of distinct astrophysical conditions and comparing our results to the latest precision mass measurements, we are able to shed light on the conditions/masses capable of producing a rare-earth peak which matches Solar data. Here we examine how our mass predictions change when using a few different sets of r-process Solar abundance residuals that have been reported in the literature. We explore how the differing error estimates of these Solar evaluations propagate through the Markov Chain Monte Carlo to our mass predictions. We find that Solar data which reports the rare-earth peak to have its highest abundance at mass number A = 162 can require distinctly different mass predictions from data with the peak centered at A = 164. Nevertheless, we find that two important general conclusions from past work, regarding the inconsistency of ‘cold’ astrophysical outflows with current mass measurements and the need for local stability at N = 104 in ‘hot’ scenarios, remain robust in the face of differing Solar data evaluations. Additionally, we show that the masses our procedure finds capable of producing a peak at A < 164 are not in line with the latest precision mass measurements.
We detail the approaches of particle swarm optimization and Bayesian inference through Markov chain Monte Carlo for equation of state development. This work includes formulation of the modeling for the equation of state, numeric optimization of the parametric models via particle swarm optimization, and generation of probability distributions of equations of state from Markov chain Monte Carlo.
Ultrasonic damage detection and characterization is commonly used in nondestructive evaluation (NDE) of aerospace composite components. In recent years there has been an increased development of guided wave based methods. In real materials and structures, these dispersive waves result in complicated behavior in the presence of complex damage scenarios. Model-based characterization methods utilize accurate three dimensional finite element models (FEMs) of guided wave interaction with realistic damage scenarios to aid in defect identification and classification. This work describes an inverse solution for realistic composite damage characterization by comparing the wavenumber-frequency spectra of experimental and simulated ultrasonic inspections. The composite laminate material properties are first verified through a Bayesian solution (Markov chain Monte Carlo), enabling uncertainty quantification surrounding the characterization. A study is undertaken to assess the efficacy of the proposed damage model and comparative metrics between the experimental and simulated output. The FEM is then parameterized with a damage model capable of describing the typical complex damage created by impact events in composites. The damage is characterized through a transdimensional Markov chain Monte Carlo solution, enabling a flexible damage model capable of adapting to the complex damage geometry investigated here. The posterior probability distributions of the individual delamination petals as well as the overall envelope of the damage site are determined.
Tristructural isotropic (TRISO) particle fuel is one of the most promising fuel concepts enabling high temperature and high burnup reactor operation. One dominant source of radioactivity released from the TRISO particles is silver (Ag), which is subject to a high release fraction and long decay life compared to other fission products. Previous modeling efforts using the fuel performance code BISON indicated nonnegligible uncertainties in modeling the diffusion process of fission products in TRISO compared to the Advanced Gas Reactor experiments. The overall uncertainties observed when modeling the fission product diffusion can result from uncertainties in model parameters, noisy experimental measurements, and deficiencies in the developed models. The three types of underlying uncertainties have not yet been properly quantified in open literature. Here, this paper presents the Bayesian uncertainty quantification (UQ) using massively parallelizable Markov chain Monte Carlo samplers. The uncertainties due to model parameters, model inadequacy, and experimental measurement noise are quantified, with the σ term used to represent the sum of the model inadequacy and measurement noise uncertainties. It is worth noting that this is the first time the σ term is inferred for nuclear fuel experiments, as compared to using prescribed values for uncertainty quantification in previous work. The parallelizable Markov chain Monte Carlo samplers efficiently infer the model parameters and the σ term, giving insight into physical parameters like diffusion coefficients and the combined model discrepancy and measurement noise. A subsequent forward uncertainty quantification (UQ) is also performed based on the calibration results to generate more accurate predictions of the Ag release. The model inadequacy plus experimental noise is the most dominant source of uncertainty compared to the parametric uncertainty. All the UQ analyses presented in this work are based on the second series of the irradiation experiments in the Advanced Gas Reactor program.
The two dimensional O(3) sigma model, just as quantum chromodynamics, is an asymptotically free theory with a mass gap. Therefore, it is an interesting and simple toy model to investigate algorithms for Markov Chain Monte Carlo simulations of quantum chromodynamics. In this talk, we discuss the construction of a trivializing map, a field transformation from a given theory to a trivial one, through a suitably chosen gradient flow. An analytic solution for the generating functional of this trivializing flow has been obtained by a perturbative expansion in the flow time. Utilizing this solution allows for new approaches to be considered when proposing updates for a Markov Chain Monte Carlo algorithm.
Inverse analysis of vibratory system is an important subject in fault identification, model updating, and robust design and control. It is challenging subject because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield results that are similar with respect to actual response measurements; and 3) uncertainties inevitably exist. The aim of this research is to leverage upon computational intelligence through statistical inference to facilitate an enhanced, probabilistic framework using incomplete modal response measurement. This new framework is built upon efficient inverse identification through optimization, whereas Bayesian inference is employed to account for the effect of uncertainties. To overcome the computational cost barrier, we adopt Markov chain Monte Carlo (MCMC) to characterize the target function/distribution. Instead of using single Markov chain in conventional Bayesian approach, we develop a new sampling theory with multiple parallel, interactive and adaptive Markov chains and incorporate it into Bayesian inference. This can harness the collective power of these Markov chains to realize the concurrent search of multiple local optima. The number of required Markov chains and their respective initial model parameters are automatically determined via Monte Carlo simulation-based sample pre-screening followed by K-means clustering analysis. These enhancements can effectively address the aforementioned challenges in finite element inverse analysis. The validity of this framework is systematically demonstrated through case studies.
As more and more nonlinear estimation techniques become available, our interest is in finding out what performance improvement, if any, they can provide for practical nonlinear problems that have been traditionally solved using linear methods. In this paper we examine the problem of estimating spacecraft position using conical scan (conscan) for NASA's Deep Space Network antennas. We show that for additive disturbances on antenna power measurement, the problem can be transformed into a linear one, and we present a general solution to this problem, with the least square solution reported in literature as a special case. We also show that for additive disturbances on antenna position, the problem is a truly nonlinear one, and we present two approximate solutions based on linearization and Unscented Transformation respectively, and one 'exact' solution based on Markov Chain Monte Carlo (MCMC) method. Simulations show that, with the amount of data collected in practice, linear methods perform almost the same as MCMC methods. It is only when we artificially reduce the amount of collected data and increase the level of noise that nonlinear methods show significantly better accuracy than that achieved by linear methods, at the expense of more computation.
Here we introduce a new training algorithm for deep neural networks that utilize random complex exponential activation functions. Our approach employs a Markov chain Monte Carlo sampling procedure to iteratively train network layers, avoiding global and gradient-based optimization while maintaining error control. It consistently attains the theoretical approximation rate for residual networks with complex exponential activation functions, determined by network complexity. Additionally, it enables efficient learning of multiscale and high-frequency features, producing interpretable parameter distributions. Despite using sinusoidal basis functions, we do not observe Gibbs phenomena in approximating discontinuous target functions.