Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian Neural Network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Image-based solar estimates

An example device is configured to determine, based on a sky image of a portion of sky over a power distribution network and using a convolutional neural network (CNN)-based image regression model, an estimated global horizontal irradiance (GHI) value and manage or control the power distribution network using the estimated GHI value. The device may also be configured to determine, based on GHI values and aggregate load values for at least a portion of the power distribution network, using a Bayesian Structural Time Series model, an estimated photovoltaic power output value for the at least a portion of the power distribution network. The device may manage or control the power distribution network using the estimated photovoltaic power output value.

Bernstein, Andrey↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deducing neutron star equation of state from telescope spectra with machine-learning-derived likelihoods

The interiors of neutron stars reach densities and temperatures beyond the limits of terrestrial experiments, providing vital laboratories for probing nuclear physics. While the star's interior is not directly observable, its pressure and density determine the star's macroscopic structure which affects the spectra observed in telescopes. The relationship between the observations and the internal state is complex and partially intractable, presenting difficulties for inference. Previous work has focused on the regression from stellar spectra of parameters describing the internal state. We demonstrate a calculation of the full likelihood of the internal state parameters given observations, accomplished by replacing intractable elements with machine learning models trained on samples of simulated stars. Our machine-learning-derived likelihood allows us to perform maximum a posteriori estimation of the parameters of interest, as well as full scans. We demonstrate the technique by inferring stellar mass and radius from an individual stellar spectrum, as well as equation of state parameters from a set of spectra. Our results are more precise than pure regression models, reducing the width of the parameter residuals by 11.8% in the most realistic scenario. The neural networks will be released as a tool for fast simulation of neutron star properties and observed spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

A Hybrid Gradient Method to Designing Bayesian Experiments for Implicit Models

Bayesian experimental design (BED) aims at designing an experiment to maximize the information gathering from the collected data. The optimal design is usually achieved by maximizing the mutual information (MI) between the data and the model parameters. When the analytical expression of the MI is unavailable, e.g.,having implicit models with intractable data distributions, a neural network-based lower bound of the MI was recently proposed and a gradient ascent method was used to maximize the lower bound [1]. However, the approach in [1] requires a pathwise sampling path to compute the gradient of the MI lower bound with respect to the design variables, and such a pathwise sampling path is usually inaccessible for implicit models. In this work, we propose a hybrid gradient approach that leverages recent advances in variational MI estimator and evolution strategies (ES)combined with black-box stochastic gradient ascent (SGA) to maximize the MI lower bound. This allows the design process to be achieved through a unified scalable procedure for implicit models without sampling path gradients. Several experiments demonstrate that our approach significantly improves the scalability of BED for implicit models in high-dimensional design space.

Zhang, Jiaxin↗

Chaotic neural dynamics facilitate probabilistic computations through sampling

Cortical neurons exhibit highly variable responses over trials and time. Theoretical works posit that this variability arises potentially from chaotic network dynamics of recurrently connected neurons. Here, we demonstrate that chaotic neural dynamics, formed through synaptic learning, allow networks to perform sensory cue integration in a sampling-based implementation. We show that the emergent chaotic dynamics provide neural substrates for generating samples not only of a static variable but also of a dynamical trajectory, where generic recurrent networks acquire these abilities with a biologically plausible learning rule through trial and error. Furthermore, the networks generalize their experience in the stimulus-evoked samples to the inference without partial or all sensory information, which suggests a computational role of spontaneous activity as a representation of the priors as well as a tractable biological computation for marginal distributions. These findings suggest that chaotic neural dynamics may serve for the brain function as a Bayesian generative model.

60 APPLIED LIFE SCIENCES↗

A flexible event reconstruction based on machine learning and likelihood principles

Event reconstruction is a central step in many particle physics experiments, turning detector observables into parameter estimates; for example estimating the energy of an interaction given the sensor readout of a detector. A corresponding likelihood function is often intractable, and approximations need to be constructed. Here, in our work, we first show how the full likelihood for a many-sensor detector can be broken apart into smaller terms, and secondly how we can train neural networks to approximate all terms solely based on forward simulation. Our technique results in a fast, flexible, and close-to-optimal surrogate model proportional to the likelihood and can be used in conjunction with standard inference techniques allowing for a consistent treatment of uncertainties. We illustrate our technique for parameter inference in neutrino telescopes based on maximum likelihood and Bayesian posterior sampling. Given its great flexibility, we also showcase our method for geometry optimization enabling to learn optimal detector designs. Lastly, we apply our method to realistic simulation of a ton-scale water-based liquid scintillator detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Optimization of Water-Alternating-CO2 Injection Field Operations Using a Machine-Learning-Assisted Workflow

Summary This paper will present a robust workflow to address multiobjective optimization (MOO) of carbon dioxide (CO2)-enhanced oil recovery (EOR)-sequestration projects with a large number of operational control parameters. Farnsworth unit (FWU) field, a mature oil reservoir undergoing CO2 alternating water injection (CO2-WAG) EOR, will be used as a field case to validate the proposed optimization protocol. The expected outcome of this work would be a repository of Pareto-optimal solutions of multiple objective functions, including oil recovery, carbon storage volume, and project economics. FWU’s numerical model is used to demonstrate the proposed optimization workflow. Because using MOO requires computationally intensive procedures, machine-learning-based proxies are introduced to substitute for the high-fidelity model, thus reducing the total computation overhead. The vector machine regression combined with the Gaussian kernel (Gaussian-SVR) is used to construct proxies. An iterative self-adjusting process prepares the training knowledge base to develop robust proxies and minimizes computational time. The proxies’ hyperparameters will be optimally designed using Bayesian optimization to achieve better generalization performance. Trained proxies will be coupled with multiobjective particle swarm Optimization (MOPSO) protocol to construct the Pareto-front solution repository. The outcomes of this workflow will be a repository containing Pareto-optimal solutions of multiple objectives considered in the CO2-WAG project. The proposed optimization workflow will be compared with another established methodology using a multilayer neural network (MLNN) to validate its feasibility in handling MOO with a large number of parameters to control. Optimization parameters used include operational variables that might be used to control the CO2-WAG process, such as the duration of the water/gas injection period, producer bottomhole pressure (BHP) control, and water injection rate of each well included in the numerical model. It is proved that the workflow coupling Gaussian-SVR proxies and the iterative self-adjusting protocol is more computationally efficient. The MOO process is made more rapid by squeezing the size of the required training knowledge base while maintaining the high accuracy of the optimized results. The outcomes of the optimization study show promising results in successfully establishing the solution repository considering multiple objective functions. Results are also verified by validating the Pareto fronts with simulation results using obtained optimized control parameters. The outcome from this work could provide field operators an opportunity to design a CO2-WAG project using as many inputs as possible from the reservoir models. The proposed work introduces a novel concept that couples Gaussian-SVR proxies with a self-adjusting protocol to increase the computational efficiency of the proposed workflow and to guarantee the high accuracy of the obtained optimized results. More importantly, the workflow can optimize a large number of control parameters used in a complex CO2-WAG process, which greatly extends its utility in solving large-scale MOO problems in various projects with similar desired outcomes.

Energy & Fuels↗

Multi-Objective Hyperparameter Optimization for Spiking Neural Network Neuroevolution

Neuroevolution has had significant success over recent years, but there has been relatively little work applying neuroevolution approaches to spiking neural networks (SNNs). SNNs are a type of neural network that includes temporal processing component, are not easily trained using other methods, and can be deployed into energy-efficient neuromorphic hardware. In this work, we investigate two evolutionary approaches for training SNNs. We explore the impact of the hyperparameters of the evolutionary approaches, including tournament size, population size, and representation type, on the performance of the algorithms. We present a multi-objective Bayesian-based hyperparameter optimization approach to tune the hyperparameters to produce the most accurate and smallest SNNs. We show that the hyperparameters can significantly affect the performance of these algorithms. We also perform sensitivity analysis and demonstrate that every hyperparameter value has the potential to perform well, assuming other hyperparameter values are set correctly.

Parsa, Maryam↗

Uncertainty quantification for Multiphase-CFD simulations of bubbly flows: a machine learning-based Bayesian approach supported by high-resolution experiments

In this paper, we developed a machine learning-based Bayesian approach to inversely quantify and reduce the uncertainties of multiphase computational fluid dynamics (MCFD) simulations for bubbly flows. The proposed approach is supported by high-resolution two-phase flow measurements, including those by double-sensor conductivity probes, high-speed imaging, and particle image velocimetry. Local distributions of key physical quantities of interest (QoIs), including the void fraction and phasic velocities, are obtained to support the Bayesian inference. In the process, the epistemic uncertainties of the closure relations are inversely quantified while the aleatory uncertainties from stochastic fluctuations of the system are evaluated based on experimental uncertainty analysis. The combined uncertainties are then propagated through the MCFD solver to obtain uncertainties of the QoIs, based on which probability-boxes are constructed for validation. The proposed approach relies on three machine learning methods: feedforward neural networks and principal component analysis for surrogate modeling, and Gaussian processes for model form uncertainty modeling. The whole process is implemented within the framework of an open-source deep learning library PyTorch with graphics processing unit (GPU) acceleration, thus ensuring the efficiency of the computation. The results demonstrate that with the support of high-resolution data, the uncertainties of MCFD simulations can be significantly reduced. The proposed approach has the potential for other applications that involve numerical models with empirical parameters.

42 ENGINEERING↗

Scalable Bayesian optimization with randomized prior networks

Several fundamental problems in science and engineering consist of global optimization tasks involving unknown high-dimensional (black-box) functions that map a set of controllable variables to the outcomes of an expensive experiment. Bayesian Optimization (BO) techniques are known to be effective in tackling global optimization problems using a relatively small number objective function evaluations, but their performance suffers when dealing with high-dimensional outputs. To overcome the major challenge of dimensionality, here we propose a deep learning framework for BO and sequential decision making based on bootstrapped ensembles of neural architectures with randomized priors. Using appropriate architecture choices, we show that the proposed framework can approximate functional relationships between design variables and quantities of interest, even in cases where the latter take values in high-dimensional vector spaces or even infinite-dimensional function spaces. In the context of BO, we augmented the proposed probabilistic surrogates with re-parameterized Monte Carlo approximations of multiple-point (parallel) acquisition functions, as well as methodological extensions for accommodating black-box constraints and multi-fidelity information sources. We test the proposed framework against state-of-the-art methods for BO and demonstrate superior performance across several challenging tasks with high-dimensional outputs, including a constrained multi-fidelity optimization task involving shape optimization of rotor blades in turbo-machinery.

97 MATHEMATICS AND COMPUTING↗

$\mathrm{SageNet}$: Fast Neural Network Emulation of the Stiff-amplified Gravitational Waves from Inflation

Accurate modeling of the inflationary gravitational waves (GWs) requires time-consuming, iterative numerical integrations of differential equations to take into account their backreaction on the expansion history. To improve computational efficiency while preserving accuracy, we present the Stiff-amplified Gravitational-wave Emulator Network (SageNet), a deep learning framework designed to replace conventional numerical solvers (code available at https://github.com/YifangLuo/SageNet). SageNet employs a long short-term memory architecture to emulate the present-day energy density spectrum of the inflationary GWs with possible stiff amplification, Ω GW (f). Trained on a data set of 25,689 numerically generated solutions, SageNet allows accurate reconstructions of Ω GW (f) and generalizes well to a wide range of cosmological parameters; 90.9% of the test emulations with randomly distributed parameters exhibit errors of under 4%. In addition, SageNet demonstrates its ability to learn and reproduce the artificial, adaptive sampling patterns in numerical calculations, which implement denser sampling of frequencies around changes in spectral indices in Ω GW (f). The dual capability of learning both physical and artificial features of the numerical GW spectra establishes SageNet as a robust alternative to exact numerical methods. Finally, our benchmark tests show that SageNet reduces the computation time from tens of seconds to milliseconds, achieving a speedup of ∼10 4 times over standard CPU-based numerical solvers with the potential for further acceleration on GPU hardware. These capabilities make SageNet a powerful tool for accelerating Bayesian inference procedures for extended cosmological models. In a broad sense, the SageNet framework offers a fast, accurate, and generalizable solution to modeling cosmological observables whose theoretical predictions demand costly differential equation solvers.

Astronomy data modeling↗

Towards a data-driven model of hadronization using normalizing flows

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing flows, termed MAGIC, that improves the agreement between simulated and experimental distributions of high-level (macroscopic) observables by adjusting single-emission (microscopic) dynamics. Our results constitute an important step toward realizing a machine-learning based model of hadronization that utilizes experimental data during training. Finally, we demonstrate how a Bayesian extension to this normalizing-flow architecture can be used to provide analysis of statistical and modeling uncertainties on the generated observable distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Extending Parsimonious Bayesian Inference

Parsimonious Bayesian inference is a theoretical framework for efficient data assimilation that seeks to balance increased consistency between predictions and training data against corresponding increases in model complexity. Within this framework, over-training is understood as optimization that encodes excessive information within model parameters while only achieving small improvements between predictions and training data. This project aims to develop practical methods of limiting excess model information during optimization. One key observation is that practical heuristics for parsimonious learning in high-dimensions must balance expressivity, i.e. the ability of the model to capture diverse predictions with only a few non-zero parameters, against discoverability, i.e. the ability to train the model with gradient-based optimization and drive parameters to low information states. As such, we developed logical activation functions that are able to adaptively approximate arbitrary truth tables that define Boolean logic operations within a probabilistic framework. These functions have demonstrated the ability to learn exclusive disjunction (XOR) and conditioned disjunction (if [condition] then [result_if_true] else [result_if_false]) within a single layer of a neural network. To efficiently exploit these activation functions to drive parsimonious learning required several other advances within the domain of variational inference. The most efficient form of complexity suppression is structured sparsification, driving most model parameters to zero while achieving the structural coherence among nonzeros needed for bandwidth reduction. Such models are not only far more efficient at suppressing information-theoretic complexity, they also reduce the other forms of complexity (computations, communication, storage, and the number of dependencies needed to evaluate predictions). Aiming to support enhanced sparsification, this project examined new approaches to high-dimensional variational inference that allow us to calibrate and control parameter uncertainty during optimization. By identifying which parameters can sustain sparsifying perturbations with little impact on prediction quality, we can develop better pruning strategies by framing them as approximate Bayesian inference. These advances also open paths to mitigate concerns with deploying advanced learning methods in resource-constrained environments, such as running models on power-limited or communication-limited devices.

97 MATHEMATICS AND COMPUTING↗

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES↗

Explainable and Differentiable Reinforcement Learning for Multi-objective Optimization in Particle Accelerators

Operating particle accelerators involves optimizing multiple goals simultaneously, which can be challenging due to trade-offs among objectives. While evolutionary algorithms like the genetic algorithm (GA) have been used for various Multi-Objective Optimization (MOO) tasks, they are not inherently suited for complex control problems. This talk highlights two variations of Reinforcement Learning (RL) for concurrently optimizing heat load and trip rates at the Continuous Electron Beam Accelerator Facility (CEBAF). The problem involves strict constraints on individual states, actions, and overall energy requirements of the beam. First, this talk highlights how differentiability can be harnessed through a Deep Differentiable Reinforcement Learning (DDRL) approach to address MOO issues within particle accelerators. We examine the DDRL method alongside Model Free Reinforcement Learning (MFRL), GA, and Bayesian Optimization (BO). The performance of these methods is assessed by generating a Pareto-front for two objectives. Our findings indicate that DDRL excels in handling high-dimensional problems more effectively than MFRL, BO, and GA. Next, we will show integration of explainable physics-based constraints into RL algorithms to enhance trans- parency and trust in decision-making processes by enabling users to verify that agents adhere to established physical principles. This surrogate function can be modeled using neural networks or sparse dictionary mod- els. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment provided but the surrogate model. In addi- tion, we find that the introduction of a mathematical functional dictionary based surrogate model enables our reinforcement learning algorithms to reliably converge for difficult high-dimensional accelerator controls environments.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Improving Subsurface Stress Characterization for Carbon Dioxide Storage Projects by Incorporating Machine Learning Techniques

The overall objective of this project is to develop a framework for reliable characterization and prediction of the state of stress in the overburden and underburden (including the basement) in CO 2 storage reservoirs using machine learning and integrated geomechanics and geophysical methods. Specifically, we propose to develop workflow encompassing of technologies and/or methods to predict stress and pressure changes due to CO 2 injection in an active tertiary recovery site and their impacts on subtle fault activation, fractures and occurrence of microseismic events and compare responses to field observations. In this project, we anticipate using dataset from the Farnsworth field Unit (FWU) which is operated by Purdure Petroleum. A novel elastic-waveform VSP inversion technique will be used to estimate high-resolution spatial and temporal changes of elastic moduli in CO 2 storage reservoirs, which will be combined with velocity-stress relationship derived from laboratory tests to obtain subsurface pressure and stress. Clustered microseismic data will be jointly inverted for improved focal mechanisms. Least-squares reverse-time migration of microseismic waveform data will be performed to directly image fracture/fault zones. Additionally, a deep neural network machine learning technique with convolutional and recurrent layers will be used for learning the spectro-temporal structures in microseismic waveforms. The results of this geotechnical data analysis will be integrated to develop a high-resolution 3D mechanical earth model extending from the overburden sealing formations to the underburden including the basement. Mechanical properties will be derived through integration of mechanical logs, tests, available results from chemo-mechanical laboratory tests, and elastic inversion of seismic data using a combination of Bayesian and stochastic methods as well as machine learning technique. Failure features (faults/fractures) will be represented and/or modeled based on seismic and core data analysis. A transient hydrodynamic-geomechanical model will be developed through coupling with the calibrated FWU reservoir simulation model. The full physics coupled model will be used to train a reduced order proxy model using machine learning algorithm for estimating stress which will then be used with appropriate constitutive relationships and forward seismological models to simulate pressure changes and induced microseismicity. An advanced optimization framework will be developed to perform a history match to minimize error between field observations and simulated. The history matched proxy model will be verified against the full-physics equivalent. The field observations that will be used in the coupled model calibration process include pressure/stress inverted from VSP, moment magnitude from microseismic analysis, real time downhole pressure measurements, production and injection data. Parameter sensitivity and uncertainty analysis will be performed to characterize the impact of model parameter uncertainty on stress estimates. The proposed project will have significant impact on future field implementation of the proposed technology. Because the project field site is an ongoing CO 2 EOR development, the value of the new technology will be demonstrated in an operational context and evaluated as a viable risk mitigation strategy. Cost/benefit will be evaluated together with the various commercial incentives for CO 2 sequestration available to oil and gas operators. The extensive available dataset and ongoing data acquisition under the SWP Phase III work plan provides flexibility for investigation of multiple approaches and reduces technical risk.

58 GEOSCIENCES↗

Training Spiking Neural Networks Using Combined Learning Approaches

Spiking neural networks (SNNs), the class of neural networks used in neuromorphic computing, are difficult to train using traditional back-propagation techniques. Spike timingdependent plasticity (STDP) is a biologically inspired learning mechanism that can be used to train SNNs. Evolutionary algorithms have also been demonstrated as a method for training SNNs. In this work, we explore the relationship between these two training methodologies. We evaluate STDP and evolutionary optimization as standalone methods for training networks, and also evaluate a combined approach where STDP weight updates are applied within an evolutionary algorithm. We also apply Bayesian hyperparameter optimization as a meta learner for each of the algorithms. We find that STDP by itself is not an ideal learning rule for randomly connected networks, while the inclusion of STDP within an evolutionary algorithm leads to similar performance, with a few interesting differences. This study suggests future work in understanding the relationship between network topology and learning rules.

Elbrecht, Daniel↗

Structure–Property Linkage in Alloys Using Graph Neural Network and Explainable Artificial Intelligence

Deep learning tools have recently shown significant potential for accelerating the prediction of microstructure–property linkage in materials. While deep neural networks like convolution neural networks (CNNs) can extract physics information from 3D microstructure images, they often require a large network architecture and substantial training time. In this research, we trained a graph neural network (GNN) using phase field generated microstructures of Ni-Al alloys to predict the evolution of mechanical properties. We found that a single GNN is capable of accurately predicting the strengthening of Ni-Al alloys with microstructures of varying sizes and dimensions, which cannot otherwise be done with a CNN. Additionally, GNN requires significantly less GPU utilization than CNN and offers more interpretable explanation of predictions using saliency analysis as features are manually defined in the graph. We also utilize explainable artificial intelligence tool Bayesian Inference to determine the coefficients in the power law equation that governs coarsening of precipitates. Overall, our work demonstrates the ability of the GNN to accurately and efficiently extract relevant information from material microstructures without having restrictions on microstructure size or dimension and offers an interpretable explanation.

Chemistry↗