Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GAUSSIAN PROCESSES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Scalable Gaussian Processes, GPyTorch Application Benchmarking, and Targeted Adaptive Design (TAD) on ThetaGPU

We aim at showcasing the scalability of Gaussian Process (GP). The naive GP implementation scales cubically with data size, which can be prohibitive, so GP has not heretofore been considered suitable for very large-scale problem settings. We take advantage of GPyTorch, a library for scalable GPs built on top of PyTorch that incorporates GPU acceleration. With GPyTorch, one can achieve nearly linear scaling with structured kernel interpolation (SKI) and constant-time predictive covariances computation with LanczOs Variance Estimates (LOVE) while preserving accuracy. We also take advantage of the computational power of ThetaGPU, a supercomputer of Argonne Leadership Computing Facility (ALCF). In addition, we implement a scalable, GPU-ready version of Targeted Adaptive Design (TAD), a GP-based data-driven algorithm that efficiently searches the control space of an advanced manufacturing experiment for settings capable of producing a required design within a specified tolerance, despite the poorly known mapping from control settings to design. We finally show our benchmarking for GPyTorch and TAD performance on CPU vs. ThetaGPU and discuss the results and implications.

97 MATHEMATICS AND COMPUTING↗

Detection Limits of Low-mass, Long-period Exoplanets Using Gaussian Processes Applied to HARPS-N Solar Radial Velocities

Radial velocity (RV) searches for Earth-mass exoplanets in the habitable zone around Sun-like stars are limited by the effects of stellar variability on the host star. In particular, suppression of convective blueshift and brightness inhomogeneities due to photospheric faculae/plage and starspots are the dominant contribution to the variability of such stellar RVs. Gaussian process (GP) regression is a powerful tool for statistically modeling these quasi-periodic variations. We investigate the limits of this technique using 800 days of RVs from the solar telescope on the High Accuracy Radial velocity Planet Searcher for the Northern hemisphere (HARPS-N) spectrograph. These data provide a well-sampled time series of stellar RV variations. Into this data set, we inject Keplerian signals with periods between 100 and 500 days and amplitudes between 0.6 and 2.4 m s{sup −1}. We use GP regression to fit the resulting RVs and determine the statistical significance of recovered periods and amplitudes. We then generate synthetic RVs with the same covariance properties as the solar data to determine a lower bound on the observational baseline necessary to detect low-mass planets in Venus-like orbits around a Sun-like star. Our simulations show that discovering planets with a larger mass (∼0.5 m s{sup −1}) using current-generation spectrographs and GP regression will require more than 12 yr of densely sampled RV observations. Furthermore, even with a perfect model of stellar variability, discovering a true exo-Venus (∼0.1 m s{sup −1}) with current instruments would take over 15 yr. Therefore, next-generation spectrographs and better models of stellar variability are required for detection of such planets.

47 OTHER INSTRUMENTATION↗

O'Hare Airport roadway traffic prediction via data fusion and Gaussian process regression

This study proposes an approach of leveraging information gathered from multiple traffic data sources at different resolutions to obtain approximate inference on the traffic distribution of Chicago's O'Hare Airport area. Specifically, it proposes the ingestion of traffic datasets at different resolutions to build spatiotemporal models for predicting the distribution of traffic volume on the road network. Due to its good adaptability and flexibility for spatiotemporal data, the Gaussian process (GP) regression was employed to provide short-term forecasts using data collected by loop detectors (sensors) and supplemented by telematics data. The GP regression is used to make predictions of the distribution of the proportion of sensor data traffic volume represented by the telematics data for each location of the sensors. Consequently, the fitted GP model can be used to determine the approximate traffic distribution for a testing location outside of the training points. Policymakers in the transportation sector can find the results of this work helpful for making informed decisions relating to current and future transportation conditions in the area.

42 ENGINEERING↗

Dynamical Mass Estimates of the β Pictoris Planetary System through Gaussian Process Stellar Activity Modeling

Nearly 15 yr of radial velocity (RV) monitoring and direct imaging enabled the detection of two giant planets orbiting the young, nearby star β Pictoris. The δ Scuti pulsations of the star, which overwhelm planetary signals, need to be carefully suppressed. In this work, we independently revisit the analysis of the RV data following a different approach than available in the literature to model the activity of the star. We show that a Gaussian process (GP) with a stochastically driven damped harmonic oscillator kernel can model the δ Scuti pulsations. It provides similar results to parametric models but with a simpler framework, using only three hyperparameters. It also enables us to model poorly sampled RV data that were excluded from previous analyses, hence extending the RV baseline by nearly five years. Altogether, the orbit and mass of both planets can be constrained from RV only, which was not possible with the parametric modeling. To characterize the system more accurately, we also perform a joint fit of all available relative astrometry and RV data. Our orbital solutions for β Pic b favor a low eccentricity of 0.029$_{−0.024}^{+0.061}$ and a relatively short period of 21.1$_{−0.8}^{+2.0}$ yr. The orbit of β Pic c is eccentric with 0.206$_{−0.063}^{+0.074}$ with a period of 3.36 ± 0.03 yr. We find model-independent masses of 11.7 ± 1.4 and 8.5 ± 0.5 M Jup for β Pic b and c, respectively, assuming coplanarity. The mass of β Pic b is consistent with the hottest start evolutionary models, at an age of 25 ± 3 Myr. A direct detection of β Pic c would provide a second calibration measurement in a coeval system.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Multi-objective Bayesian alloy design using multi-task Gaussian processes

In design applications, correlations among material properties (such as the tendency for stronger materials to be less ductile) are often neglected. This approach is echoed in multi-objective optimization techniques which treat each performance characteristic as an independent objective, aiming to optimize scalar functions and find optimal Pareto fronts. However, this overlooks the statistical relationships between performance characteristics inherent in a material system. To address this, we propose the use of Bayesian optimization, a highly efficient black-box optimization algorithm known for constructing Gaussian processes (GPs) – uncorrelated surrogates - to model objective functions. Rather than evaluating multiple GPs for each objective function separately, we argue for a shift towards jointly modeling these objective functions, considering their statistical correlations. This integrated approach utilizes naturally occurring relationships among material properties, providing additional information to enhance the performance of the design framework. This requires the replacement of multiple independent GPs with a single multi-task GP, employing a correlation matrix to construct a multi-task kernel function, wherein each task corresponds to a single objective function. Here, we anticipate this refined methodology will better leverage material correlations, improving design optimization results.

36 MATERIALS SCIENCE↗

Physics makes the difference: Bayesian optimization and active learning via augmented Gaussian process

Abstract Both experimental and computational methods for the exploration of structure, functionality, and properties of materials often necessitate the search across broad parameter spaces to discover optimal experimental conditions and regions of interest in the image space or parameter space of computational models. The direct grid search of the parameter space tends to be extremely time-consuming, leading to the development of strategies balancing exploration of unknown parameter spaces and exploitation towards required performance metrics. However, classical Bayesian optimization (BO) strategies based on the Gaussian process (GP) do not readily allow for the incorporation of the known physical behaviors or past knowledge. Here we explore a hybrid optimization/exploration algorithm created by augmenting the standard GP with a structured probabilistic model of the expected system’s behavior. This approach balances the flexibility of the non-parametric GP approach with a rigid structure of physical knowledge encoded into the parametric model. The fully Bayesian treatment of the latter allows additional control over the optimization via the selection of priors for the model parameters. The method is demonstrated for a noisy version of a standard univariate test function used to evaluate optimization algorithms and further extended to physical lattice models. This methodology is expected to be universally suitable for injecting prior knowledge in the form of physical models and past data in the BO framework.

42 ENGINEERING↗

Physics-constrained Gaussian process model for prediction of hydrodynamic interactions between wave energy converters in an array

To improve the efficiency of wave farms and achieve maximum power generation, the layout of wave energy converters (WECs) in an array needs to be carefully designed so that the hydrodynamic interactions can be positively exploited. For this, the hydrodynamic characteristics of the WEC array in different layouts need to be calculated. However, such calculations using numerical models usually entail significant computational cost, especially for large arrays of WECs. To address the computational challenge, a physics-constrained Gaussian process (GP) model is proposed to replace the original expensive numerical model and predict the hydrodynamic characteristics of the WECs for any array layout. By exploring the relationship between the WEC array (i.e., the input) and different hydrodynamic characteristics (i.e., the output), here we summarize a set of physical constraints/features, including invariance, symmetry, and additivity. This prior knowledge about the input-output relationship is then directly embedded in the constructed GP model through the design of physics-constrained kernels. In particular, a double-sum invariant kernel is first developed to incorporate the invariance and symmetry features, and then an additive kernel is developed to incorporate the additive feature of the problem. The invariant kernel and the additive kernel are then integrated to construct the physics-constrained GP model. Compared to the standard GP model, the proposed physics-constrained GP models require less training data to achieve the desired accuracy in predicting the hydrodynamic characteristics and are also less vulnerable to the curse of dimensionality (i.e., good scalability for large arrays) due to the use of an additive kernel. The efficiency, accuracy, and scalability of the proposed approach are demonstrated through an application to predict the hydrodynamic characteristics for WEC arrays of different sizes and layouts.

16 TIDAL AND WAVE POWER↗

HostSub_GP: Precise Galaxy Background Subtraction in Transient Long-slit Spectroscopy with Gaussian Processes

We present a novel host galaxy subtraction technique in long-slit spectroscopy for extragalactic transients. Unlike classic methods which generally estimate the background using simple interpolation of local galaxy flux in the 2D spectrum, our approach leverages multi-band archival images of the host galaxies to model the background emission from the galaxy in the 2D spectrum. Such imaging encodes the wavelength-dependent galaxy profile along the slit, and is readily accessible through wide-field imaging surveys. We construct a smooth prior for the 2D galaxy profile with a Gaussian process (GP) based on these reference images, and use another GP to model the correlated deviations from the prior in the observed spectrum. This enables accurate inference of the galaxy flux blended with the transient. On synthetic long-slit data of a spiral galaxy extracted from a Multi Unit Spectroscopic Explorer hyper-spectral cube, the GP method remains robust as long as the host galaxy is spatially resolved and consistently outperforms classic methods. We apply the method to archival Keck spectra of two real transients, SN 2019eix and AT 2019qiz, to further demonstrate how the method uniquely recovers weak spectral features amid strong galaxy contamination, enabling refined constraints on the properties of both transients. We have released the software implementation, HostSub_GP, a scalable toolkit that leverages JAX, with an MIT license.

79 ASTRONOMY AND ASTROPHYSICS↗

Estimation of the Ambient Wind Field From Wind Turbine Measurements Using Gaussian Process Regression

In the search for a lower levelized cost of wind energy, one approach is to increase the accuracy of wind turbine measurements such as wind speed and wind direction. The sensors available on wind turbines are susceptible to local turbulence and measurement bias, which can result in suboptimal turbine performance. As an alternative, recent research has considered using the sensor measurements in a coordinated manner. With such a cooperative approach, the local wind conditions can be estimated more accurately and reliably without the need for additional measurement equipment. In this paper, a novel wind field estimation approach is presented that estimates the local wind conditions based on turbine measurements using Gaussian processes. We show that the estimation framework is able to improve the accuracy of the wind direction estimate both in an offline and online manner, as well as identify possible biases in the sensors and reduce unnecessary wind turbine yaw activity.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

A Gaussian process based surrogate approach for the optimization of cylindrical targets

Simulating direct-drive inertial confinement experiments presents significant computational challenges, both due to the complexity of the codes required for such simulations and the substantial computational expense associated with target design studies. Machine learning models, and in particular, surrogate models, offer a solution by replacing simulation results with a simplified approximation. In this study, we apply surrogate modeling and optimization techniques that are well established in the existing literature to one-dimensional simulation data of a new cylindrical target design containing deuterium–tritium fuel. These models predict yields without the need for expensive simulations. We find that Bayesian optimization with Gaussian process surrogates enhances sampling efficiency in low-dimensional design spaces but becomes less efficient as dimensionality increases. Nonetheless, optimization routines within two-dimensional and five-dimensional design spaces can identify designs that maximize yield, while also aligning with established physical intuition. Optimization routines, which ignore constraints on hydrodynamic instability growth, are shown to lead to unstable designs in 2D, resulting in yield loss. However, routines that utilize 1D simulations and impose constraints on the in-flight aspect ratio converge on novel cylindrical target designs that are stable against hydrodynamic instability growth in 2D and achieve high yield.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Estimation of hydraulic conductivity in a watershed using sparse multi-source data via Gaussian process regression and Bayesian experimental design

Enhanced water management systems depend on accurate estimation of subsurface hydraulic properties. However, geologic formations can vary significantly, so information from a single source (e.g., widely spaced boreholes) is insufficient in characterizing subsurface aquifer properties. Therefore, multiple sources of information are needed to complement the hydrogeology understanding of a region. Here, this study presents a numerical framework in which information from different measurement sources is combined to characterize the 3D random field in a multi-fidelity prediction model. Coupled with the model, a Bayesian experimental design was used to determine the best future sampling locations. The Upper Sangamon watershed in east-central Illinois was selected as the case study site, where the multi-fidelity Gaussian process model was used to estimate the hydraulic conductivity in the region of interest. Multi-source observation data were obtained from electrical resistivity and borehole pumping tests. The accuracy of the model prediction is dependent on the locations and the distribution of both high- and low-fidelity data. Furthermore, the multi-fidelity model was compared with the single-fidelity model. The uncertainties and confidence in the measurements and parameter estimates were quantified and used to design future cycles of data collection to further improve the confidence intervals.

54 ENVIRONMENTAL SCIENCES↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

GPLaSDI: Gaussian Process Latent Space Dynamics Identification

GPLaSDI is a software that introduces GPLaSDI, a novel LaSDI-based framework that relies on Gaussian process (GP) for latent space ODE interpolations. Using GPs offers two significant advantages. First, it enables the quantification of uncertainty over the reduced order model (ROM) predictions. Second, leveraging this prediction uncertainty allows for efficient adaptive training through a greedy selection of additional training data points. This approach does not require prior knowledge of the underlying PDEs. Consequently, GPLaSDI is inherently non-intrusive and can be applied to problems without a known PDE or its residual. We demonstrate the effectiveness of our approach on the Burgers equation, Vlasov equation for plasma physics, and a rising thermal bubble problem. Our proposed method achieves between 200 and 100,000 times speed-up, with up to 7% relative error.

Choi, Youngsoo↗

Proxy quality control of biomass particles using thermogravimetric analysis and Gaussian process regression models

Abstract The temperature experienced by reactants during preparation in a reactor is a key component in determining the yield and homogeneity of usable chemical products such as biomass particles. Thermocouples with sensors can be used to monitor spatial temperature gradients within reactors but these sensors are often too expensive and/or invasive. The present work proposes a strategy to identify optimal machine learning models to infer the maximum effective temperature experienced by particles during oxidative biomass torrefaction using key thermochemical combustion parameters. The maximum rate of weight loss, the corresponding temperature, and fixed carbon content on a dry‐ash‐free basis are used as literature‐based predictor variables obtained from thermogravimetric analysis. The evaluation of 24 machine‐learning models using the standard tenfold cross‐validation method suggests that the exponential Gaussian process regression (GPR) model is the most effective, followed by other GPR models. These high‐performing GPR models were also utilized to predict the effective preparation temperature distribution of reactor‐produced biomass particles under eight conditions of varying residence time and air‐to‐biomass ratio. The effective preparation temperature and residence time of individual biomass particles were then encoded into the torrefaction severity factor and used to estimate the energy yield of the reactor output as a novel quality control method. © 2023 The Authors. Biofuels, Bioproducts and Biorefining published by Society of Industrial Chemistry and John Wiley & Sons Ltd.

09 BIOMASS FUELS↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗