Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep Gaussian processes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Exploring the energy landscape of RBMs: reciprocal space insights into bosons, hierarchical learning and symmetry breaking

Deep generative models have become ubiquitous due to their ability to learn and sample from complex distributions. Despite the proliferation of various frameworks, the relationships among these models remain largely unexplored, a gap that hinders the development of a unified theory of AI learning. In this work, we address two central challenges: clarifying the connections between different deep generative models and deepening our understanding of their learning mechanisms. We focus on Restricted Boltzmann Machines (RBMs), a class of generative models known for their universal approximation capabilities for discrete distributions. By introducing a reciprocal space formulation for RBMs, we reveal a connection between these models, diffusion processes, and systems of coupled bosons. Our analysis shows that at initialization, the RBM operates at a saddle point, where the local curvature is determined by the singular values of the weight matrix, whose distribution follows the Marc̆enko-Pastur law and exhibits rotational symmetry. During training, this rotational symmetry is broken due to hierarchical learning, where different degrees of freedom progressively capture features at multiple levels of abstraction. This leads to a symmetry breaking in the energy landscape, reminiscent of Landau’s theory. This symmetry breaking in the energy landscape is characterized by the singular values and the weight matrix eigenvector matrix. We derive the corresponding free energy in a mean-field approximation. We show that in the limit of infinite size RBM, the reciprocal variables are Gaussian distributed. Our findings indicate that in this regime, there will be some modes for which the diffusion process will not converge to the Boltzmann distribution. To illustrate our results, we trained replicas of RBMs with different hidden layer sizes using the MNIST dataset. Our findings not only bridge the gap between disparate generative frameworks but also shed light on the fundamental processes underpinning learning in deep generative models.

97 MATHEMATICS AND COMPUTING↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Adaptive Deep Reinforcement Learning Algorithm for Distribution System Cyber Attack Defense With High Penetration of DERs

With grid modernization, smart inverters are increasingly used to execute advanced controls for distribution network reliability. However, this also increases the cyber-attack space. Here this paper focuses on the defense approaches to restore the system to normal operation circumstances in the presence of cyber-attacks. A unique deep reinforcement learning (DRL) method is developed to minimize voltage violations and reduce power losses for impacted feeders. The defense problem is reformulated as a Markov decision-making process to dynamically control DERs while minimizing load shedding. This is achieved via an improved soft actor-critic (SAC)-based DRL algorithm, which can govern DER set points and load-shedding scenarios in discrete and continuous modes via the auto-tune entropy and Gaussian policy features. Numerical comparison results on the modified IEEE 123-node system with other control approaches, such as Volt-VAR (VV), Volt-Watt (VW), and model predictive control (MPC) show that the proposed method can eliminate voltage violations and provide feasible control actions that perform complete mitigation of cyber-threats.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep learning categorization of infrasound array data

Here we develop a deep learning-based infrasonic detection and categorization methodology that uses convolutional neural networks with self-attention layers to identify stationary and non-stationary signals in infrasound array processing results. Using features extracted from the coherence and direction-of-arrival information from beamforming at different infrasound arrays, our model more reliably detects signals compared with raw waveform data. Using three infrasound stations maintained as part of the International Monitoring System, we construct an analyst-reviewed data set for model training and evaluation. We construct models using a 4-category framework, a generalized noise vs non-noise detection scheme, and a signal-of-interest (SOI) categorization framework that merges short duration stationary and non-stationary categories into a single SOI category. We evaluate these models using a combination of k-fold cross-validation, comparison with an existing “state-of-the-art” detector, and a transportability analysis. Although results are mixed in distinguishing stationary and non-stationary short duration signals, f-scores for the noise vs non-noise and SOI analyses are consistently above 0.96, implying that deep learning-based infrasonic categorization is a highly accurate means of identifying signals-of-interest in infrasonic data records.

47 OTHER INSTRUMENTATION↗

The Nonradiative Properties of Self–Trapped Holes in Ultra–Wide Bandgap Gallium Oxide Film

The photoluminescence (PL) of self-trapped holes (STH) in ultra-wide bandgap β-Ga 2 O 3 is commonly its most dominant light emission and is an inherent property. Thus, gaining knowledge of the crystal dynamics that impact the PL properties is vital to sensor and other technologies. The PL, Raman-phonons, and their interactions are studied at an extreme temperature range of 77–622 K. The PL is studied up to the bandgap value of ≈5 eV. It is found that the high-energy Raman modes provide a major route to the nonradiative process of the PL via STH–phonon interaction with an activation energy of 72 meV. This dynamic is modeled with the configurational coordinate scheme at the strong phonon coupling limit. The exceptionally broad Gaussian PL linewidth manifests this coupling. The weak temperature response of the PL energy peak position indicates that the STH has characteristics of a deep-level defect. This contrasts with the large redshift of ≈220 meV of the optical gap of the film, ascertained from transmission. Unlike the temperature response of the high-energy phonons, the behavior of the low-energy phonons is found to follow the Bose–Einstein population increase, indicating no strong interaction with the STH.

Raman↗

MOOSE ProbML: Parallelized probabilistic machine learning and uncertainty quantification for computational energy applications

Here, this paper presents the development and demonstration of massively parallel probabilistic machine learning (ML) and uncertainty quantification (UQ) capabilities within the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source computational platform for parallel finite element and finite volume analyses. In addressing the computational expense and uncertainties inherent in complex multiphysics simulations, this paper integrates Gaussian process (GP) variants, active learning, Bayesian inverse UQ, adaptive forward UQ, Bayesian optimization, evolutionary optimization, and Markov chain Monte Carlo (MCMC) within MOOSE. It also elaborates on the interaction among key MOOSE systems — Sampler, MultiApp, Reporter, and Surrogate — in enabling these capabilities. The modularity offered by these systems enables development of a multitude of probabilistic ML and UQ algorithms in MOOSE. Example code demonstrations include parallel active learning and parallel Bayesian inference via active learning. The impact of these developments is illustrated through five applications relevant to computational energy applications: UQ of nuclear fuel fission product release, using parallel active learning Bayesian inference; very rare events analysis in nuclear microreactors using active learning; advanced manufacturing process modeling using multi-output GPs (MOGPs) and dimensionality reduction; fluid flow using deep GPs (DGPs); and tritium transport model parameter optimization for fusion energy, using batch Bayesian optimization. These capabilities are part of the MOOSE framework.

97 - MATHEMATICS AND COMPUTING↗

Combining High-Throughput Experiments and Active Learning to Characterize Deep Eutectic Solvents

The high tunability of deep eutectic solvents (DESs) stems from the ease of changing their precursors and relative compositions. However, measuring the physicochemical properties across large composition and temperature ranges, necessary to properly design target-specific DESs, is tedious and error-prone and represents a bottleneck in the advancement and scalability of DES-based applications. As such, active learning (AL) methodologies based on Gaussian processes (GPs) were developed in this work to minimize the experimental effort necessary to characterize DESs. Owing to its importance for large-scale applications, the reduction of DES viscosity through the addition of a low-molecular-weight solvent was explored as a case study. A high-throughput experimental screening was initially performed on nine different ternary DESs. Then, GPs were successfully trained to predict DES viscosity from its composition and temperature, showcasing the ability of these stochastic, nonparametric models to accurately describe the physicochemical properties of complex mixtures. Finally, the ability of GPs to provide estimates of their own uncertainty was leveraged through an AL framework to minimize the number of data points necessary to obtain accurate viscosity modes. This led to a significant reduction in data requirements, with many systems requiring only five independent viscosity data points to be properly described.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automatic parameter selection for electron ptychography via Bayesian optimization

Abstract Electron ptychography provides new opportunities to resolve atomic structures with deep sub-angstrom spatial resolution and to study electron-beam sensitive materials with high dose efficiency. In practice, obtaining accurate ptychography images requires simultaneously optimizing multiple parameters that are often selected based on trial-and-error, resulting in low-throughput experiments and preventing wider adoption. Here, we develop an automatic parameter selection framework to circumvent this problem using Bayesian optimization with Gaussian processes. With minimal prior knowledge, the workflow efficiently produces ptychographic reconstructions that are superior to those processed by experienced experts. The method also facilitates better experimental designs by exploring optimized experimental parameters from simulated data.

42 ENGINEERING↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

A robust estimator of mutual information for deep learning interpretability

Abstract We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning (DL) models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced ‘Jimmie’), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established MI estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train DL models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available in this GitHub repository.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Bayesian high-throughput estimation of plasma temperature and density from emission spectroscopy

Here, this paper introduces a novel approach for automated high-throughput estimation of plasma temperature and density using atomic emission spectroscopy, integrating Bayesian inference with sophisticated physical models. We provide an in-depth examination of Bayesian methods applied to the complexities of plasma diagnostics, supported by a robust framework of physical and measurement models. Our methodology is demonstrated using experimental observations in the field of magneto-inertial fusion, focusing on individual and sequential shot analyses of the Plasma Liner Experiment at LANL. The results demonstrate the effectiveness of our approach in enhancing the accuracy and reliability of plasma parameter estimation and in using the analysis to reveal the deep hidden structure in the data. This study not only offers a new perspective of plasma analysis but also paves the way for further research and applications in nuclear instrumentation and related domains.

Bayesian inference↗

Rough surface scattering results based on bandpass autocorrelation forms

Surface-height autocorrelation forms such as Gaussian and exponential are often used in studies of near-normal incidence rough-surface scattering. Such models require the existence of a constant, or DC, value in the spectrum. The consequences of autocorrelation forms that correspond to spectral processes that are essentially bandpass in nature are examined. One such process is that of ocean wind waves. In this case, the spectral components do not extend down to zero frequency. The physical optics backscatter theory is reexamined relative to such autocorrelation functions. Experimental results obtained from a wavetank are compared to the autocorrelation model used in the analysis. The analysis indicates that Gaussian correlation length or mean-square slope is not an appropriate parameter for narrowband conditions and that significant slope is a more relevant parameter. Inherent in the deep-phase assumption is some form of slope dependency. The analysis given (and variants thereof) can be used to provide insight into the physical effects of separate spectral components and of spectral directionality.

Miller, Lee S.↗

A Decision-Making Machine Learning Approach in Hermite Spectral Approximations of Partial Differential Equations

The accuracy and effectiveness of Hermite spectral methods for the numerical discretization of partial differential equations on unbounded domains are strongly affected by the amplitude of the Gaussian weight function employed to describe the approximation space. This is particularly true if the problem is under-resolved, i.e., there are no enough degrees of freedom. The issue becomes even more crucial when the equation under study is time-dependent, forcing in this way the choice of Hermite functions where the corresponding weight depends on time. In order to adapt dynamically the approximation space, it is here proposed an automatic decision-making process that relies on machine learning techniques, such as deep neural networks and support vector machines. The algorithm is numerically tested with success on a simple 1D problem, but the main goal is its exportability in the context of more serious applications. Here we also show at the end an application in the framework of plasma physics.

97 MATHEMATICS AND COMPUTING↗

Effects of correlated noise on the full-spectrum combining and complex-symbol combining arraying techniques

The process of combining telemetry signals received at multiple antennas, commonly referred to as arraying, can be used to improve communication link performance in the Deep Space Network (DSN). By coherently adding telemetry from multiple receiving sites, arraying produces an enhancement in signal-to-noise ratio (SNR) over that achievable with any single antenna in the array. A number of different techniques for arraying have been proposed and their performances analyzed in past literature. These analyses have compared different arraying schemes under the assumption that the signals contain additive white Gaussian noise (AWGN) and that the noise observed at distinct antennas is independent. In situations where an unwanted background body is visible to multiple antennas in the array, however, the assumption of independent noises is no longer applicable. A planet with significant radiation emissions in the frequency band of interest can be one such source of correlated noise. For example, during much of Galileo's tour of Jupiter, the planet will contribute significantly to the total system noise at various ground stations. This article analyzes the effects of correlated noise on two arraying schemes currently being considered for DSN applications: full-spectrum combining (FSC) and complex-symbol combining (CSC). A framework is presented for characterizing the correlated noise based on physical parameters, and the impact of the noise correlation on the array performance is assessed for each scheme.

Vazirani, P.↗

Quantum Simulation of Molecular Dynamics Processes─A Benchmark Study Using a Classical Simulator and Present-Day Quantum Hardware

Here, we explore how the fundamental problems in quantum molecular dynamics can be modeled using classical simulators (emulators) of quantum computers and the actual quantum hardware available to us today. The list of problems we tackle includes propagation of a free wave packet, vibration of a harmonic oscillator, and tunneling through a barrier. Each of these problems starts with the initial wave packet setup. Although Qiskit provides a general method for initializing wave functions, in most cases it generates deep quantum circuits. While these circuits perform well on noiseless simulators, they suffer from excessive noise on quantum hardware. To overcome this issue, we designed a shallower quantum circuit for preparing a Gaussian-like initial wave packet, which improves the performance of real hardware. Next, quantum circuits are implemented to apply the kinetic and potential energy operators for the evolution of a wave function over time. The results of our modeling on classical emulators of quantum hardware agree perfectly with the results obtained using the traditional (classical) methods. This serves as a benchmark and demonstrates that the quantum algorithms and Qiskit codes we developed are accurate. However, the results obtained on the actual quantum hardware available today, such as IBM’s superconducting qubits and IonQ’s trapped ions, indicate large discrepancies due to hardware limitations. This work highlights both the potential and challenges of using quantum computers to solve fundamental quantum molecular dynamics problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving Dark Energy Constraints Using Low-Redshift Large-Scale Structures

The primary goal of this project was to improve constraints on dark energy measurements by improving our ability to extract cosmological information from low redshift large-scale structures. PI Clowe's project's primary aim was to reduce the bias in measurements of the masses of clusters of galaxies to a level where the evolution of the cluster mass function can be used in the Vera Rubin Observatory's Legacy Survey of Space and Time Dark Energy Science Collaboration survey to improve the accuracy of the measurement of dark energy and other cosmological parameters. Co-PI Seo's project studied observational systematics affecting large-scale clustering of galaxies, which will be used to improve dark energy constraints from the Dark Energy Spectroscopic Instrument (DESI). The cluster lensing project employed a series of simulations and observations of clusters of galaxies to test numerous potential systematic errors in cluster mass measurements using weak gravitational lensing as the accuracy of current weak lensing measurements are more than order of magnitude worse than what is required to use clusters of galaxies for accurate determination of dark energy parameters. PI Clowe and group developed and analyzed simulations to test for, and correct biases introduced in, the weak lensing measurement process. Finally, PI Clowe and group developed a method of detecting clusters using galaxy overdensities and applied the method to the BLISS and DES surveys. The success of spectroscopic dark energy mission such as the extended Baryon Oscillation Spectroscopic Survey (eBOSS) and the Dark Energy Spectroscopic Instrument (DESI) will depend on a thorough understanding of various observational systematics in the target density fluctuations that would give rise to spurious, non-cosmological signals. PI Seo and group developed a deep learning, artificial neural network (ANN) technique that modeled and mitigated such effects, aimed at deriving more robust galaxy clustering signals not only for the baryon acoustic oscillation feature and redshift-space distortions but also for primordial non-Gaussianity constraint.

79 ASTRONOMY AND ASTROPHYSICS↗

Discriminative Dimensionality Reduction using Deep Neural Networks for Clustering of LIGO Data

In this paper, leveraging the capabilities of neural networks for modeling the non-linearities that exist in the data, we propose several models that can project data into a low dimensional, discriminative, and smooth manifold. The proposed models can transfer knowledge from the domain of known classes to a new domain where the classes are unknown. A clustering algorithm is further applied in the new domain to find potentially new classes from the pool of unlabeled data. The research problem and data for this paper originated from the Gravity Spy project which is a side project of Advanced Laser Interferometer Gravitational-wave Observatory (LIGO). The LIGO project aims at detecting cosmic gravitational waves using huge detectors. However non-cosmic, non-Gaussian disturbances known as "glitches", show up in gravitational-wave data of LIGO. This is undesirable as it creates problems for the gravitational wave detection process. Gravity Spy aids in glitch identification with the purpose of understanding their origin. Since new types of glitches appear over time, one of the objective of Gravity Spy is to create new glitch classes. Towards this task, we offer a methodology in this paper to accomplish this.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Surrogate Modeling of Nonlinear Dynamic Systems: A Comparative Study

Surrogate models play a vital role in overcoming the computational challenge in designing and analyzing nonlinear dynamic systems, especially in the presence of uncertainty. This paper presents a comparative study of different surrogate modeling techniques for nonlinear dynamic systems. Four surrogate modeling methods, namely, Gaussian process (GP) regression, a long short-term memory (LSTM) network, a convolutional neural network (CNN) with LSTM (CNN-LSTM), and a CNN with bidirectional LSTM (CNN-BLSTM), are studied and compared. All these model types can predict the future behavior of dynamic systems over long periods based on training data from relatively short periods. The multi-dimensional inputs of surrogate models are organized in a nonlinear autoregressive exogenous model (NARX) scheme to enable recursive prediction over long periods, where current predictions replace inputs from the previous time window. Three numerical examples, including one mathematical example and two nonlinear engineering analysis models, are used to compare the performance of the four surrogate modeling techniques. The results show that the GP-NARX surrogate model tends to have more stable performance than the other three deep learning (DL)-based methods for the three particular examples studied. The tuning effort of GP-NARX is also much lower than its deep learning-based counterparts.

42 ENGINEERING↗