Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian process fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Robustness of the Stochastic Parameterization of Subgrid-Scale Wind Variability in Sea Surface Fluxes

Abstract High-resolution numerical models have been used to develop statistical models of the enhancement of sea surface fluxes resulting from spatial variability of sea surface wind. In particular, studies have shown that flux enhancement is not a deterministic function of the resolved state. Previous studies focused on single geographical areas or used a single high-resolution numerical model. This study extends the development of such statistical models by considering six different high-resolution models, four different geographical regions, and three different 10-day periods, allowing for a systematic investigation of the robustness of both the deterministic and stochastic parts of the data-driven parameterization. Results indicate that the deterministic part, based on regressing the unresolved normalized flux onto resolved-scale normalized flux and precipitation, is broadly robust across different models, regions, and time periods. The statistical features of the stochastic part of the model (spatial and temporal autocorrelation and parameters of a Gaussian process fit to the regression residual) are also found to be robust and not strongly sensitive to the underlying model, modeled geographical region, or time period studied. Best-fit Gaussian process parameters display robust spatial heterogeneity across models, indicating potential for improvements to the statistical model. These results illustrate the potential for the development of a generic, explicitly stochastic parameterization of sea surface flux enhancements dependent on wind variability.

Endo, Kota↗

Code for multi-shape Gaussian process (GP) fitting with uncertainty quantification (UQ)

This is code associated with the publication “Nonparametric Multi-shape Modeling with Uncertainty Quantification,” authored by Hengrui Luo (Lawrence Berkeley National Laboratory) and Justin Strait (Los Alamos National Laboratory). The code is used to fit multiple-output Gaussian process (GP) models of planar closed curves to collections of ordered point sets, allowing for flexible nonlinear prediction of the underlying curve under dense or sparse point set samplings and with or without noise, as well as tractable uncertainty quantification. To do this, we employ use of a periodic kernel to account for the nonlinear input space of closed curves, and combine with coregionalization models to account for dependence both (i) between curve coordinates, and (ii) between pairs of curves. Functions in the code are capable of fitting these models, as well as performing additional tasks with the fitted curves such as (a) shape registration and alignment, (b) shape averaging, and (c) fitting for curve sub-populations / clusters.

Strait, Justin↗

Bright Opportunities for Atmospheric Characterization of Small Planets: Masses and Radii of K2-3 b, c, and d and GJ3470 b from Radial Velocity Measurements and Spitzer Transits

We report improved masses, radii, and densities for four planets in two bright M-dwarf systems, K2-3 and GJ3470, derived from a combination of new radial velocity and transit observations. Supplementing K2 photometry with follow-up Spitzer transit observations refined the transit ephemerides of K2-3 b, c, and d by over a factor of 10. We analyze ground-based photometry from the Evryscope and Fairborn Observatory to determine the characteristic stellar activity timescales for our Gaussian Process fit, including the stellar rotation period and activity region decay timescale. The stellar rotation signals for both stars are evident in the radial velocity data and is included in our fit using a Gaussian process trained on the photometry. We find the masses of K2-3 b, K2-3 c, and GJ3470 b to be 6.48(+0.99,-0.93), 2.14(+1.08,-1.28), and 12.58(+1.31,-1.28)Mꚛ, respectively. K2-3 d was not significantly detected and has a 3σ upper limit of 2.80 Mꚛ. These two systems are training cases for future TESS systems; due to the low planet densities (ρ < 3.7 g/cu.cm) and bright host stars (K < 9 mag), they are among the best candidates for transmission spectroscopy in order to characterize the atmospheric compositions of small planets.

Molly R. Kosiarek↗

Bayesian D‐Optimal Designs for Gaussian Process Surrogate Models

Computer experiments often employ space-filling strategies to create surrogate models with strong predictive performance. The impact of model parameter estimation for Gaussian process surrogates, however, is often overlooked. Obtaining a better initial estimate of the covariance lengthscale parameter, θ, can greatly improve the resulting Gaussian process fit through more effective sequential acquisitions during active learning. In this work, we propose a novel initial design maximizing the Bayesian D-optimality criterion of the Gaussian process lengthscale parameter. Previously published results have shown the emphasis on lengthscale estimation to be promising, but relied on an empirically driven design creation process. Our Bayesian D-optimal designs are rooted in information theory and lead to more informative sequential acquisitions by improving lengthscale estimation. In many cases, these gains eventually result in better surrogates than those seeded with space-filling initial designs. Furthermore, Bayesian D-optimal designs can be tailored to either isotropic or anisotropic covariance structures, and the Bayesian framework enables the inclusion of prior knowledge in the design process, offering greater flexibility and adaptability. Through several simulation studies, we demonstrate the advantages of Bayesian D-optimal designs in terms of both lengthscale estimation accuracy and predictive performance during active learning.

Bayesian experimental design↗

Sequential Bayesian Methods for Analyzing Computer Models

Efficient analysis of computer models is essential for the validation and uncertainty quantification of those models. Surrogate models, and Gaussian processes in particular, are a common and powerful approach to analyzing computer models that treat computer models as a black-box function. Gaussian processes form a Bayesian model over a space of functions that gives a measure of uncertainty about the computer model output at unobserved locations and a framework for sequential sampling. This tutorial will show how to fit Gaussian processes on a series of test functions and apply sequential design techniques to estimate extrema, level sets, and reliabilities of those test functions.

97 MATHEMATICS AND COMPUTING↗

Considerations for Optimizing the Photometric Classification of Supernovae from the Rubin Observatory

The Vera C. Rubin Observatory will increase the number of observed supernovae (SNe) by an order of magnitude; however, it is impossible to spectroscopically confirm the class for all SNe discovered. Thus, photometric classification is crucial, but its accuracy depends on the not-yet-finalized observing strategy of Rubin Observatory's Legacy Survey of Space and Time (LSST). We quantitatively analyze the impact of the LSST observing strategy on SNe classification using simulated multiband light curves from the Photometric LSST Astronomical Time-Series Classification Challenge (PLAsTiCC). First, we augment the simulated training set to be representative of the photometric redshift distribution per SNe class, the cadence of observations, and the flux uncertainty distribution of the test set. Then we build a classifier using the photometric transient classification library snmachine, based on wavelet features obtained from Gaussian process fits, yielding a similar performance to the winning PLAsTiCC entry. We study the classification performance for SNe with different properties within a single simulated observing strategy. We find that season length is important, with light curves of 150 days yielding the highest performance. Cadence also has an important impact on SNe classification; events with median inter-night gap <3.5 days yield higher classification performance. Interestingly, we find that large gaps (>10 days) in light-curve observations do not impact performance if sufficient observations are available on either side, due to the effectiveness of the Gaussian process interpolation. This analysis is the first exploration of the impact of observing strategy on photometric SN classification with LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

dfdjaxGP

A small python package to fit Gaussian processes using Jax and leveraging the automatic differentiation in Jax to predict arbitrary derivatives from the GP. This package is meant to supplement the scientific community use of Gaussian process prediction with derivatives. The package is designed to smoothly work standalone or be used with the numpyro probabilistic programming language.

Grosskopf, Micheal↗

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES↗

A method for predicting failure statistics for steady state elevated temperature structural components

This paper presents the initial development of a high temperature life prediction method that accounts for the variability in the material properties of Grade 91 steel. The method accounts for material variability by fitting a variable 3-parameter Weibull distribution to experimental rupture data and accounts for the variability of creep deformation on the steady-state stresses via a Monte Carlo approach. To ensure reasonable computational times, the model represents the material as an extremely viscous Stokes fluid with a non-Newtonian viscosity, therefore solving the stress relaxation problem with a steady, static, instead of transient, analysis. Furthermore, the complete statistical analysis combines this model for creep deformation with a probabilistic model for creep rupture to evaluate the probability of premature failure for a set of sample problems, comparing the predicted failure statistics to the design life predicted by the ASME Boiler and Pressure Vessel Code rules.

42 ENGINEERING↗

Adaptive Sensing of Time Series with Application to Remote Exploration

We address the problem of adaptive informationoptimal data collection in time series. Here a remote sensor or explorer agent throttles its sampling rate in order to track anomalous events while obeying constraints on time and power. This problem is challenging because the agent has limited visibility -- all collected datapoints lie in the past, but its resource allocation decisions require predicting far into the future. Our solution is to continually fit a Gaussian process model to the latest data and optimize the sampling plan on line to maximize information gain. We compare the performance characteristics of stationary and nonstationary Gaussian process models. We also describe an application based on geologic analysis during planetary rover exploration. Here adaptive sampling can improve coverage of localized anomalies and potentially benefit mission science yield of long autonomous traverses.

artificial intelligence↗

Data-Efficient Methods for Determining Flory–Huggins χ Parameters in Multicomponent Polymer Formulations

Polymer formulations are essential in diverse applications including personal care products, coatings, paints, adhesives, and plastic materials. Designing these formulations requires navigating large, complex design spaces, where phase and self-assembly behavior critically impact performance. The Flory–Huggins χ parameter, which quantifies segmental miscibility, is widely used to parametrize the excess free energy of mixing in formulation models. In this work, we introduce two data-efficient, top-down methods for estimating χ parameters using the Random Phase Approximation (RPA): (i) Boundary Nonlinear Regression (Boundary-NLR), which fits theoretical spinodal boundaries to experimental phase boundaries, and (ii) Surrogate Model Inverse Parameter Estimation (SMIPE), which uses a Gaussian Process Classifier to fit sparse phase maps via a surrogate model. Both methods allow rapid parametrization of polymer field-theoretic models without the need for additional experiments. We evaluate these approaches on data sets involving polymer–solvent–nonsolvent ternary mixtures and block copolymer–solvent systems, demonstrating their robustness to experimental noise and their relevance for real-world formulation design.

copolymers↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improving the astrometric solution of the Hyper Suprime-Cam with anisotropic Gaussian processes

Context. We study astrometric residuals from a simultaneous fit of Hyper Suprime-Cam images. Aims. We aim to characterize these residuals and study the extent to which they are dominated by atmospheric contributions for bright sources. Methods. We used Gaussian process interpolation with a correlation function (kernel) measured from the data to smooth and correct the observed astrometric residual field. Results. We find that a Gaussian process interpolation with a von Kármán kernel allows us to reduce the covariances of astrometric residuals for nearby sources by about one order of magnitude, from 30 mas 2 to 3 mas 2 at angular scales of ~1 arcmin. This also allows us to halve the rms residuals. Those reductions using Gaussian process interpolation are similar to recent result published with the Dark Energy Survey dataset. We are then able to detect the small static astrometric residuals due to the Hyper Suprime-Cam sensors effects. We discuss how the Gaussian process interpolation of astrometric residuals impacts galaxy shape measurements, particularly in the context of cosmic shear analyses at the Rubin Observatory Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantitative estimation of granitoid composition from thermal infrared multispectral scanner (TIMS) data, Desolation Wilderness, northern Sierra Nevada, California

We have produced images that quantitatively depict modal and chemical parameters of granitoids using an image processing algorithm called MINMAP that fits Gaussian curves to normalized emittance spectra recovered from thermal infrared multispectral scanner (TIMS) radiance data. We applied the algorithm to TIMS data from the Desolation Wilderness, an extensively glaciated area near the northern end of the Sierra Nevada batholith that is underlain by Jurassic and Cretaceous plutons that range from diorite and anorthosite to leucogranite. The wavelength corresponding to the calculated emittance minimum lambda(sub min) varies linearly with quartz content, SiO2, and other modal and chemical parameters. Thematic maps of quartz and silica content derived from lambda(sub min) values distinguish bodies of diorite from surrounding granite, identify outcrops of anorthosite, and separate felsic, intermediate, and mafic rocks.

Sabine, Charles↗

Bayesian Active Learning for Scanning Probe Microscopy: From Gaussian Processes to Hypothesis Learning

Recent progress in machine learning methods and the emerging availability of programmable interfaces for scanning probe microscopes (SPMs) have propelled automated and autonomous microscopies to the forefront of attention of the scientific community. However, enabling automated microscopy requires the development of task-specific machine learning methods, understanding the interplay between physics discovery and machine learning, and fully defined discovery workflows. This, in turn, requires balancing the physical intuition and prior knowledge of the domain scientist with rewards that define experimental goals and machine learning algorithms that can translate these to specific experimental protocols. Here, we discuss the basic principles of Bayesian active learning and illustrate its applications for SPM. We progress from the Gaussian process as a simple data-driven method and Bayesian inference for physical models as an extension of physics-based functional fits to more complex deep kernel learning methods, structured Gaussian processes, and hypothesis learning. These frameworks allow for the use of prior data, the discovery of specific functionalities as encoded in spectral data, and exploration of physical laws manifesting during the experiment. Here, the discussed framework can be universally applied to all techniques combining imaging and spectroscopy, SPM methods, nanoindentation, electron microscopy and spectroscopy, and chemical imaging methods and can be particularly impactful for destructive or irreversible measurements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

OGLE-2017-BLG-1186: First Application of Asteroseismology and Gaussian Processes to microlensing

We present the analysis of the event OGLE-2017-BLG-1186 from the 2017 Spitzer microlensing campaign. This is a remarkable microlensing event because its source is photometrically bright and variable, which makes it possible to perform an asteroseismic analysis using ground-based data. We find that the source star is an oscillating red giant with average timescale of ∼9 d. The asteroseismic analysis also provides us source properties including the source angular size (∼27 μas) and distance (∼11.5 kpc), which are essential for inferring the properties of the lens. When fitting the light curve, we test the feasibility of Gaussian processes (GPs) in handling the correlated noise caused by the variable source. We find that the parameters from the GP model are generally more loosely constrained than those from the traditional χ(exp 2) minimization method. We note that this event is the first microlensing system for which asteroseismology and GPs have been used to account for the variable source. With both finite-source effect and microlens parallax measured, we find that the lens is likely a ∼0.045 Mʘ brown dwarf at distance ∼9.0 kpc, or a ∼0.073 Mʘ ultracool dwarf at distance ∼9.8 kpc. Combining the estimated lens properties with a Bayesian analysis using a Galactic model, we find a ∼ 35 per cent probability for the lens to be a bulge object and ∼ 65 per cent to be a background disc object.

S.-S. Li↗

A localized ensemble of approximate Gaussian processes for fast sequential emulation

More attention has been given to the computational cost associated with the fitting of an emulator. Substantially less attention is given to the computational cost of using that emulator for prediction. This is primarily because the cost of fitting an emulator is usually far greater than that of obtaining a single prediction, and predictions can often be obtained in parallel. In many settings, especially those requiring Markov Chain Monte Carlo, predictions may arrive sequentially and parallelization is not possible. In this case, using an emulator procedure which can produce accurate predictions efficiently can lead to substantial time savings in practice. In this paper, we propose a global model approximate Gaussian process framework via extension of a popular local approximate Gaussian process (laGP) framework. Our proposed emulator can be viewed as a treed Gaussian process where the leaf nodes are laGP models, and the tree structure is learned greedily as a function of the prediction stream. The suggested method (called leapGP) has interpretable tuning parameters which control the time‐memory trade‐off. One reasonable choice of settings leads to an emulator with a training cost and makes predictions rapidly with an asymptotic amortized cost of .

97 MATHEMATICS AND COMPUTING↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗