Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep Gaussian process regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

Robust Deep Gaussian Process-Based Probabilistic Electrical Load Forecasting Against Anomalous Events

The abnormal events, such as the unprecedented COVID-19 pandemic, can significantly change the load behaviors, leading to huge challenges for traditional short-term forecasting methods. This article proposes a robust deep Gaussian processes (DGP)-based probabilistic load forecasting method using a limited number of data. Since the proposed method only requires a limited number of training samples for load forecasting, it allows us to deal with extreme scenarios that cause short-term load behavior changes. In particular, the load forecasting at the beginning of abnormal event is cast as a regression problem with limited training samples and solved by double stochastic variational inference DGP. The mobility data are also utilized to deal with the uncertainties and pattern changes and enhance the flexibility of the forecasting model. The proposed method can quantify the uncertainties of load forecasting outcomes, which would be essential under uncertain inputs. Extensive comparison results with other state-of-the-art point and probabilistic forecasting methods show that our proposed approach can achieve high forecasting accuracies with only a limited number of data while maintaining the excellent performance of capturing the forecasting uncertainties.

anomalous events↗

Decentralized Voltage Control of Large-Scale Distribution System with PVs Based on MADRL

This paper proposes a model-free decentralized control framework for the voltage regulation of large-scale distribution systems through the coordinated control of PV inverters. This is achieved by developing a novel interaction mechanism between the surrogate model and the centralized training and decentralized execution multiagent deep reinforcement learning framework. Specifically, the sparse Gaussian processes regression method is first utilized to develop the surrogate model of the original distribution system for reward calculation during the training stage, where each agent represents a sub-region in the centralized fashion for coordination strategy learning. After that, the learned control rules are used to inform controllers within each sub-region for real-time decisions with only local measurements. Comparative tests among various methods on the EPRI Ckt5 test system demonstrate the effectiveness of the proposed method.

distribution system↗

Designing Monte Carlo Simulation and an Optimal Machine Learning to Optimize and Model Space Missions

This paper investigates applying artificial intelligence (AI) algorithms to attitude control system of satellites to optimally tune the controller using high performance computing. This methodology is applied to the Virtual Telescope for X-ray Observation mission, which is a precise formation of two separate spacecraft observing multiple objects in the space in the X-ray domain. The mission is divided into phases based on the instrumentation and the mission goal. To reach an stable precise formation robust to stochastic slew and slew rate (i.e., Euler angles and angular velocities) in a minimal constrained time T , consumed energy of the attitude control system, denoted as E, and root-mean-square state error of attitude control system, denoted as e, are minimized. Monte-Carlo simulation is used for the sensitivity analysis of optimization and designing a controller. Deep neural networks (DNN), Gaussian processes (GP), and support vector regression (SVR) learn this optimization as a surrogate model, while their hyperparameters are optimized in a novel approach. THETA supercomputer at Argonne Leadership Computing Facility (ALCF) is used for optimizing the hyperparameters of DNN. The surrogate model meets the requirements of the mission, and it shows a better performance over the optimization and Monte-Carlo. The optimal DNN can satisfy the mission requirements e and T while reducing E for 90% compared to the other given methods.

42 ENGINEERING↗

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

Uncertainty Aware Deep Learning for Particle Accelerators

Standard deep learning models for classification and regression applications are ideal for capturing complex system dynamics. However, their predictions can be arbitrarily inaccurate when the input samples are not similar to the training data. Implementation of distance aware uncertainty estimation can be used to detect these scenarios and provide a level of confidence associated with their predictions. In this paper, we present results from using Deep Gaussian Process Approximation (DGPA) methods for errant beam prediction at Spallation Neutron Source (SNS) accelerator (classification) and we provide an uncertainty aware surrogate model for the Fermi National Accelerator Lab (FNAL) Booster Accelerator Complex (regression).

Rajput, Kishansingh↗

Reducing Ground-based Astrometric Errors with Gaia and Gaussian Processes

Stochastic field distortions caused by atmospheric turbulence are a fundamental limitation to the astrometric accuracy of ground-based imaging. This distortion field is measurable at the locations of stars with accurate positions provided by the Gaia DR2 catalog; we develop the use of Gaussian process regression (GPR) to interpolate the distortion field to arbitrary locations in each exposure. We introduce an extension to standard GPR techniques that exploits the knowledge that the 2D distortion field is curl-free. Applied to several hundred 90 s exposures from the Dark Energy Survey as a test bed, we find that the GPR correction reduces the variance of the turbulent astrometric distortions ≈12× , on average, with better performance in denser regions of the Gaia catalog. The rms per-coordinate distortion in the riz bands is typically ≈7 mas before any correction and ≈2 mas after application of the GPR model. The GPR astrometric corrections are validated by the observation that their use reduces, from 10 to 5 mas rms, the residuals to an orbit fit to riz-band observations over 5 yr of the r = 18.5 trans-Neptunian object Eris. We also propose a GPR method, not yet implemented, for simultaneously estimating the turbulence fields and the 5D stellar solutions in a stack of overlapping exposures, which should yield further turbulence reductions in future deep surveys.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Learning functional priors and posteriors from data and physics

In this work, we develop a new Bayesian framework based on deep neural networks to be able to extrapolate in space-time using historical data and to quantify uncertainties arising from both noisy and gappy data in physical problems. Specifically, the proposed approach has two stages: (1) prior learning and (2) posterior estimation. At the first stage, we employ the physics-informed Generative Adversarial Networks (PI-GAN) to learn a functional prior either from a prescribed function distribution, e.g., Gaussian process, or from historical data and physics. At the second stage, we employ the Hamiltonian Monte Carlo (HMC) method to estimate the posterior in the latent space of PI-GANs. In addition, we use two different approaches to encode the physics: (1) automatic differentiation, used in the physicsinformed neural networks (PINNs) for scenarios with explicitly known partial differential equations (PDEs), and (2) operator regression using the deep operator network (DeepONet) for PDE-agnostic scenarios. We then test the proposed method for (1) meta-learning for one-dimensional regression, and forward/inverse PDE problems (combined with PINNs); (2) PDE-agnostic physical problems (combined with DeepONet), e.g., fractional diffusion as well as saturated stochastic (100-dimensional) flows in heterogeneous porous media; and (3) spatial-temporal regression problems, i.e., inference of a marine riser displacement field using experimental data from the Norwegian Deepwater Programme (NDP). The results demonstrate that the proposed approach can provide accurate predictions as well as uncertainty quantification given very limited scattered and noisy data, since historical data could be available to provide informative priors. In summary, the proposed method is capable of learning flexible functional priors, e.g., both Gaussian and non-Gaussian process, and can be readily extended to big data problems by enabling mini-batch training using stochastic HMC or normalizing flows since the latent space is generally characterized as low dimensional.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Correspondence of NNGP Kernel and the Matérn Kernel

Kernels representing limiting cases of neural network architectures have recently gained popularity. However, the application and performance of these new kernels compared to existing options, such as the Matérn kernel, is not well studied. We take a practical approach to explore the neural network Gaussian process (NNGP) kernel and its application to data in Gaussian process regression. We first demonstrate the necessity of normalization to produce valid NNGP kernels and explore related numerical challenges. We further demonstrate that the predictions from this model are quite inflexible, and therefore do not vary much over the valid hyperparameter sets. We then demonstrate a surprising result that the predictions given from the NNGP kernel correspond closely to those given by the Matérn kernel under specific circumstances, which suggests a deep similarity between overparameterized deep neural networks and the Matérn kernel. Finally, we demonstrate the performance of the NNGP kernel as compared to the Matérn kernel on three benchmark data cases, and we conclude that for its flexibility and practical performance, the Matérn kernel is preferred to the novel NNGP in practical applications.

97 MATHEMATICS AND COMPUTING↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Automated Bayesian high-throughput estimation of plasma temperature and density from emission spectroscopy

Here, this paper introduces a novel approach for automated high-throughput estimation of plasma temperature and density using atomic emission spectroscopy, integrating Bayesian inference with sophisticated physical models. We provide an in-depth examination of Bayesian methods applied to the complexities of plasma diagnostics, supported by a robust framework of physical and measurement models. Our methodology is demonstrated using experimental observations in the field of magneto-inertial fusion, focusing on individual and sequential shot analyses of the Plasma Liner Experiment at LANL. The results demonstrate the effectiveness of our approach in enhancing the accuracy and reliability of plasma parameter estimation and in using the analysis to reveal the deep hidden structure in the data. This study not only offers a new perspective of plasma analysis but also paves the way for further research and applications in nuclear instrumentation and related domains.

Bayesian inference↗

Surrogate Modeling of Nonlinear Dynamic Systems: A Comparative Study

Surrogate models play a vital role in overcoming the computational challenge in designing and analyzing nonlinear dynamic systems, especially in the presence of uncertainty. This paper presents a comparative study of different surrogate modeling techniques for nonlinear dynamic systems. Four surrogate modeling methods, namely, Gaussian process (GP) regression, a long short-term memory (LSTM) network, a convolutional neural network (CNN) with LSTM (CNN-LSTM), and a CNN with bidirectional LSTM (CNN-BLSTM), are studied and compared. All these model types can predict the future behavior of dynamic systems over long periods based on training data from relatively short periods. The multi-dimensional inputs of surrogate models are organized in a nonlinear autoregressive exogenous model (NARX) scheme to enable recursive prediction over long periods, where current predictions replace inputs from the previous time window. Three numerical examples, including one mathematical example and two nonlinear engineering analysis models, are used to compare the performance of the four surrogate modeling techniques. The results show that the GP-NARX surrogate model tends to have more stable performance than the other three deep learning (DL)-based methods for the three particular examples studied. The tuning effort of GP-NARX is also much lower than its deep learning-based counterparts.

42 ENGINEERING↗

Peri-Net-Pro: the neural processes with quantified uncertainty for crack patterns

Abstract This paper develops a deep learning tool based on neural processes (NPs) called the Peri-Net-Pro, to predict the crack patterns in a moving disk and classifies them according to the classification modes with quantified uncertainties. In particular, image classification and regression studies are conducted by means of convolutional neural networks (CNNs) and NPs. First, the amount and quality of the data are enhanced by using peridynamics to theoretically compensate for the problems of the finite element method (FEM) in generating crack pattern images. Second, case studies are conducted with the prototype microelastic brittle (PMB), linear peridynamic solid (LPS), and viscoelastic solid (VES) models obtained by using the peridynamic theory. The case studies are performed to classify the images by using CNNs and determine the suitability of the PMB, LBS, and VES models. Finally, a regression analysis is performed on the crack pattern images with NPs to predict the crack patterns. The regression analysis results confirm that the variance decreases when the number of epochs increases by using the NPs. The training results gradually improve, and the variance ranges decrease to less than 0.035. The main finding of this study is that the NPs enable accurate predictions, even with missing or insufficient training data. The results demonstrate that if the context points are set to the 10th, 100th, 300th, and 784th, the training information is deliberately omitted for the context points of the 10th, 100th, and 300th, and the predictions are different when the context points are significantly lower. However, the comparison of the results of the 100th and 784th context points shows that the predicted results are similar because of the Gaussian processes in the NPs. Therefore, if the NPs are employed for training, the missing information of the training data can be supplemented to predict the results.

Mathematics↗