Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian Neural Network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Uncertainty quantification for deep learning in particle accelerator applications

With the advent of increased computational resources and improved algorithms, machine learning-based models are being increasingly applied to complex problems in particle accelerators. However, such data-driven models may provide overly confident predictions with unknown errors and uncertainties. For reliable deployment of machine learning models in high-regret and safety-critical systems such as particle accelerators, estimates of prediction uncertainty are needed along with accurate point predictions. In this investigation, we evaluate Bayesian neural networks (BNN) as an approach that can provide accurate predictions along with reliably quantified uncertainties for particle accelerator problems, and compare their performance with bootstrapped ensembles of neural networks. We select three accelerator setups for this evaluation: a storage ring, a photoinjector, and a linac. The problems span different data volumes and dimensionalities (e.g., scalar predictions as well as image outputs). It is found that BNN provide accurate predictions of the mean along with reliable estimates of predictive uncertainty across the test cases. In this vein, BNN may offer an attractive alternative to deterministic deep learning tools to generate accurate predictions with quantified uncertainties in particle accelerator applications.

43 PARTICLE ACCELERATORS↗

mmbo (Multi-modal Bayesian Optimization) [SWR-25-119]

This package implements Bayesian Neural Network (BNN) based surrogate models for multi-modal data. These models are fit using a custom Variational Bayes procedure leveraging conjugate posteriors in the last layer of the network, offering more accurate predictions with better-calibrated uncertainty quantification than mean-field Variational Bayes.

Taylor, Ian [National Laboratory of the Rockies (N↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Iterative HOMER with uncertainties

We present iHOMER, an iterative version of the HOMER method to extract Lund fragmentation functions from experimental data. Through iterations, we address the information gap between latent and observable phase spaces and systematically remove bias. To quantify uncertainties on the inferred weights, we use a combination of Bayesian neural networks and uncertainty-aware regression. We find that the combination of iterations and uncertainty quantification produces well-calibrated weights that accurately reproduce the data distribution. A parametric closure test shows that the iteratively learned fragmentation function is compatible with the true fragmentation function.

Butter, Anja [Heidelberg Univ. (Germany); Sorbonne↗

Mapping Stochastic Devices to Probabilistic Algorithms

Probabilistic and Bayesian neural networks have long been proposed as a method to incorporate uncertainty about the world (both in training data and operation) into artificial intelligence applications. One approach to making a neural network probabilistic is to leverage a Monte Carlo sampling approach that samples a trained network while incorporating noise. Such sampling approaches for neural networks have not been extensively studied due to the prohibitive requirement of many computationally expensive samples. While the development of future microelectronics platforms that make this sampling more efficient is an attractive option, it has not been immediately clear how to sample a neural network and what the quality of random number generation should be. This research aimed to start addressing these two fundamental questions by examining basic “off the shelf” neural networks can be sampled through a few different mechanisms (including synapse “dropout” and neuron “dropout”) and examine how these sampling approaches can be evaluated both in terms of evaluating algorithm effectiveness and the required quality of random numbers.

97 MATHEMATICS AND COMPUTING↗

Probabilistic Nanomagnetic Memories for Uncertain and Robust Machine Learning

This project evaluated the use of emerging spintronic memory devices for robust and efficient variational inference schemes. Variational inference (VI) schemes, which constrain the distribution for each weight to be a Gaussian distribution with a mean and standard deviation, are a tractable method for calculating posterior distributions of weights in a Bayesian neural network such that this neural network can also be trained using the powerful backpropagation algorithm. Our project focuses on domain-wall magnetic tunnel junctions (DW-MTJs), a powerful multi-functional spintronic synapse design that can achieve low power switching while also opening the pathway towards repeatable, analog operation using fabricated notches. Our initial efforts to employ DW-MTJs as an all-in-one stochastic synapse with both a mean and standard deviation didn’t end up meeting the quality metrics for hardware-friendly VI. In the future, new device stacks and methods for expressive anisotropy modification may make this idea still possible. However, as a fall back that immediately satisfies our requirements, we invented and detailed how the combination of a DW-MTJ synapse encoding the mean and a probabilistic Bayes-MTJ device, programmed via a ferroelectric or ionically modifiable layer, can robustly and expressively implement VI. This design includes a physics-informed small circuit model, that was scaled up to perform and demonstrate rigorous uncertainty quantification applications, up to and including small convolutional networks on a grayscale image classification task, and larger (Residual) networks implementing multi-channel image classification. Lastly, as these results and ideas all depend upon the idea of an inference application where weights (spintronic memory states) remain non-volatile, the retention of these synapses for the notched case was further interrogated. These investigations revealed and emphasized the importance of both notch geometry and anisotropy modification in order to further enhance the endurance of written spintronic states. In the near future, these results will be mapped to effective predictions for room temperature and elevated operation DW-MTJ memory retention, and experimentally verified when devices become available.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Comparing Automated Posterior Estimation Techniques for Modeling Strong Lenses In Ground-based Survey Data

Current and future ground-based cosmological surveys, such as the Dark Energy Survey (DES), and the Vera Rubin Observatory Legacy Survey of Space and Time (LSST), are predicted to discover thousands to tens of thousands of strong gravitational lenses. The large number of strong lenses discoverable in future surveys will make strong lensing a highly competitive and complementary cosmic probe. However, conventional lens modeling techniques are unable to scale up to the sheer number of lenses that will be discovered through upcoming surveys. Therefore, the use of automated lens analysis techniques is necessary. We demonstrate that machine learning methods can be used to automate the inference of informative model posteriors of strong lensing systems in ground-based surveys with credible uncertainty estimation. We present two Simulation-Based Inference (SBI) approaches for lens parameter estimation of galaxy-galaxy lenses. We demonstrate applications of Neural Posteriors Estima tors (NPEs) and Bayesian Neural Network (BNNs) to automate the inference of a 12-parameter lensing system for DES-like ground-based imaging data. We apply a suite of diagnostics (e.g., posterior coverage and SBC) to validate the performance of our methods. We find that NPEs outperform the BNN, producing posterior distributions that are for the most part both more accurate and more precise; in particular, several source-light model parameters are systematically biased in the BNN implementation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Comparative Analysis of Machine Learning Models for Day-Ahead Photovoltaic Power Production Forecasting

A main challenge for integrating the intermittent photovoltaic (PV) power generation remains the accuracy of day-ahead forecasts and the establishment of robust performing methods. The purpose of this work is to address these technological challenges by evaluating the day-ahead PV production forecasting performance of different machine learning models under different supervised learning regimes and minimal input features. Specifically, the day-ahead forecasting capability of Bayesian neural network (BNN), support vector regression (SVR), and regression tree (RT) models was investigated by employing the same dataset for training and performance verification, thus enabling a valid comparison. The training regime analysis demonstrated that the performance of the investigated models was strongly dependent on the timeframe of the train set, training data sequence, and application of irradiance condition filters. Furthermore, accurate results were obtained utilizing only the measured power output and other calculated parameters for training. Consequently, useful information is provided for establishing a robust day-ahead forecasting methodology that utilizes calculated input parameters and an optimal supervised learning approach. Finally, the obtained results demonstrated that the optimally constructed BNN outperformed all other machine learning models achieving forecasting accuracies lower than 5%.

14 SOLAR ENERGY↗

Global Framework for Emulation of Nuclear Calculations

We introduce a hierarchical framework that combines ab initio many-body calculations with a Bayesian neural network, developing emulators capable of accurately predicting nuclear properties across isotopic chains simultaneously and being applicable to different regions of the nuclear chart. We benchmark our developments using the oxygen isotopic chain, achieving accurate results for ground-state energies and nuclear charge radii, while providing robust uncertainty quantification. Our framework enables global sensitivity analysis of nuclear binding energies and charge radii with respect to the low-energy constants that describe the nuclear force.

FOS: Computer and information sciences↗

Strong Lensing Parameter Estimation on Ground-Based Imaging Data Using Simulation-Based Inference

Current ground-based cosmological surveys, such as the Dark Energy Survey (DES), are predicted to discover thousands of galaxy-scale strong lenses, while future surveys, such as the Vera Rubin Observatory Legacy Survey of Space and Time (LSST) will increase that number by 1-2 orders of magnitude. The large number of strong lenses discoverable in future surveys will make strong lensing a highly competitive and complementary cosmic probe. To leverage the increased statistical power of the lenses that will be discovered through upcoming surveys, automated lens analysis techniques are necessary. We present two Simulation-Based Inference (SBI) approaches for lens parameter estimation of galaxy-galaxy lenses. We demonstrate the successful application of Neural Posterior Estimation (NPE) to automate the inference of a 12-parameter lens mass model for DES-like ground-based imaging data. We compare our NPE constraints to a Bayesian Neural Network (BNN) and find that it outperforms the BNN, producing posterior distributions that are for the most part both more accurate and more precise; in particular, several source-light model parameters are systematically biased in the BNN implementation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Bayesian Inference with Latent Hamiltonian Neural Networks (L-HNNs)

When sampling for Bayesian inference, one popular approach is to use Hamiltonian Monte Carlo (HMC) and the No-U-Turn Sampler (NUTS). However, HMC and NUTS can require numerous numerical gradients of the target density and can prove slow in practice. We propose Hamiltonian neural networks (HNNs) with HMC and NUTS for solving Bayesian inference problems [1, 2]. Once trained, HNNs do not require gradients of the target density while sampling. Moreover, they satisfy important properties such as perfect time reversibility and Hamiltonian conservation, making them well suited for use within HMC and NUTS because stationarity can be shown. We also propose an HNN extension called latent HNNs (L-HNNs), which predict latent variable outputs. Compared to HNNs, L-HNNs offer improved expressivity and a reduction in integration errors. Finally, we propose employing L-HNNs in NUTS with an online error monitoring scheme to prevent degeneracy of the sampling in regions of low probability density. We demonstrate L-HNNs in NUTS with online error monitoring by using several example cases involving complex, heavy-tailed, and high local curvature probability densities. Overall, L-HNNs in NUTS with online error monitoring satisfactorily inferred these probability densities. Compared to traditional NUTS, L-HNNs in NUTS with online error monitoring improved the effective sample size (ESS) per gradient by an order of magnitude.

97 MATHEMATICS AND COMPUTING↗

Efficient Bayesian inference with latent Hamiltonian neural networks in No-U-Turn Sampling

When sampling for Bayesian inference, one popular approach in the computational field is to use Hamiltonian Monte Carlo (HMC) and specifically the No-U-Turn Sampler (NUTS), which automatically decides the end time of the Hamiltonian trajectory. However, HMC and NUTS can require numerous numerical gradients of the target density and can prove slow in practice when relying on computationally expensive forward models. We propose Latent Hamiltonian neural networks (L-HNNs) with HMC and NUTS for solving Bayesian inference problems. Once trained, L-HNNs do not require numerical gradients of the target density during sampling, and hence numerous evaluations of the forward computational model. Moreover, L-HNNs satisfy important properties such as perfect time reversibility and Hamiltonian conservation, making them well-suited for use within HMC and NUTS because stationarity can be shown. We also propose the integration of L-HNNs in an online error monitoring scheme, in which numerical gradients of the target density are used for a few samples whenever the L-HNNs prediction errors are large. This online error monitor scheme prevents sample degeneracy in regions of low probability density and ensures robust uncertainty quantification. We demonstrate L-HNNs in NUTS with online error monitoring on several analytical examples involving complex, heavy-tailed, and high-local-curvature probability densities. We then demonstrate the applicability of L-HNNs in NUTS to two computational case studies, namely the Allen-Cahn stochastic partial differential equation and an elliptic partial differential equation with 25 and 50 inference parameters, respectively. Overall, the L-HNNs in NUTS with online error monitoring satisfactorily inferred these probability densities. In conclusion, compared to traditional NUTS, L-HNNs in NUTS with online error monitoring required 1–2 orders of magnitude fewer numerical gradients of the target density and improved the effective sample size (ESS) per gradient (which is a measure of both the sampling quality and the computational expense) by an order of magnitude.

97 MATHEMATICS AND COMPUTING↗

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES↗

Improved naive Bayesian probability classifier in predictions of nuclear mass

Recently, novel statistical methods such as neural networks and Bayesian learning methods are implemented to describe the nuclear masses. Based on previous studies, an improved naive Bayesian probability (iNBP) classifier is proposed to study the nuclear masses by refining the results of sophisticated nuclear models. In the iNBP method, the prediction for nuclear masses is treated as a classification problem. The residuals are classified into several groups to generate prior and conditional probabilities, and the posterior probabilities are further determined by the Bayesian formula. We choose the expectation with maximum probability as the final prediction. Reliability of the iNBP method is assessed by analyzing the global optimizations and the extrapolating capabilities. Here, the iNBP method exhibits impressive improvements on global descriptions for different mass models. Moreover, the method shows robust extrapolating capabilities. Results demonstrate the iNBP method can be applied to predict the nuclear masses of unknown regions. Considering the local mass relations, the iNBP method can offer considerable fine-tuning of the mass descriptions from nuclear models. The methodology proposed in this paper can also be applied to other model-based extrapolations of nuclear observables.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Missing data in multi-omics integration: Recent advances through artificial intelligence

Biological systems function through complex interactions between various ‘omics (biomolecules), and a more complete understanding of these systems is only possible through an integrated, multi-omic perspective. This has presented the need for the development of integration approaches that are able to capture the complex, often non-linear, interactions that define these biological systems and are adapted to the challenges of combining the heterogenous data across ‘omic views. A principal challenge to multi-omic integration is missing data because all biomolecules are not measured in all samples. Due to either cost, instrument sensitivity, or other experimental factors, data for a biological sample may be missing for one or more ‘omic techologies. Recent methodological developments in artificial intelligence and statistical learning have greatly facilitated the analyses of multi-omics data, however many of these techniques assume access to completely observed data. A subset of these methods incorporate mechanisms for handling partially observed samples, and these methods are the focus of this review. We describe recently developed approaches, noting their primary use cases and highlighting each method's approach to handling missing data. We additionally provide an overview of the more traditional missing data workflows and their limitations; and we discuss potential avenues for further developments as well as how the missing data issue and its current solutions may generalize beyond the multi-omics context.

97 MATHEMATICS AND COMPUTING↗