Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian process fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The Mira-Titan Universe. III. Emulation of the Halo Mass Function

We construct an emulator for the halo mass function over group and cluster mass scales for a range of cosmologies, including the effects of dynamical dark energy and massive neutrinos. The emulator is based on the recently completed Mira-Titan Universe suite of cosmological N-body simulations. The main set of simulations spans 111 cosmological models with 2.1 Gpc boxes. We extract halo catalogs in the redshift range z = [0.0, 2.0] and for masses M-200c >= 10(13)M(circle dot)/h. The emulator covers an eight-dimensional hypercube spanned by {Omega(m)h(2), Omega(b)h(2), Omega(nu)h(2), sigma(8), h, n(s), w(0), w(a)}; spatial flatness is assumed. We obtain smooth halo mass functions by fitting piecewise second-order polynomials to the halo catalogs and employ Gaussian process regression to construct the emulator while keeping track of the statistical noise in the input halo catalogs and uncertainties in the regression process. For redshifts z less than or similar to 1, the typical emulator precision is better than 2% for 10(13)-10(14)M(circle dot)/h and <10<^> M similar or equal to 101(circle dot)(54)/h. For comparison, fitting functions using the traditional universal form for the halo mass function can be biased at up to 30% at M similar or equal to 10(144)M(circle dot)/h for z = 0. Our emulator is publicly available at https://github.com/SebastianBocquet/MiraTitanHMFemulator.

cosmology: theory↗

A Gaussian Process Enhancement to Linear Parameter Varying Models

Simulation and analysis for modern engineering systems now routinely requires the merging of multiple disciplines, physical-domains, time-scales, and data sets — all at ever increasing levels. These capabilities are especially needed in the domain of Advanced Air Mobility, where rapidly emerging vehicle designs are significantly more complex, while having to be both cost-effective and safe. To meet these engineering challenges, machine learning methods are an attractive option for merging models and data across multiple areas while providing uncertainty quantification and maintaining computational efficiency. This paper examines the use of Gaussian process machine learning to generalize and enhance the commonly used class of quasi-Linear Parameter Varying models for fast full-envelope simulation while also supporting control system design and analysis with model uncertainty. Gaussian process machine learning is selected because it: can fuse multiple data sets, enables an easy trade-off between data fitting and smoothing, provides model uncertainty quantification, scales well with increasing complexity, and does not generally require starting from a large training data set. To demonstrate the benefits of the approach, a robust stability analysis with Gaussian process uncertainty is shown for a NASA reference design of an electric quad-rotor air-taxi concept vehicle with motor parameter uncertainty.

Gaussian Process↗

Simulation of Inflated Pahoehoe Lava Flows

A new stochastic model simulates late-stage pahoehoe lobes where random processes dominate emplacement. The model prescribes probabilistic rules for determining where and when parcels of lava move within the lobe. Unlike a classical Brownian motion random walk, the model allows individual parcels to remain dormant, but fluid, for multiple time steps. The randomness of parcel volume transfers within the lobe interior as well as at the margins qualitatively reflects inflation processes observed in the field. The fraction of inflated volume to total volume increases with the total volume, with greater than 75% of the lobe volume contributed through inflation for typical lobes. The influence on planform shape and topographic cross-sectional profiles of total volume, source area and shape, topographic confinement, and sequential breakouts at the lobe margins, are all explored with the stochastic model. Each of these factors influences the overall lobe thickness and width. The model provides a means for assessing the relative importance of these processes through comparisons with field data. For the first time, Gaussian and parabolic functions are quantitatively fit to field measurements of pahoehoe lobes. Both functional forms provide adequate description of the cross-sectional flow shapes. When comparing simulated lobes to field data, sequential breakouts at the lobe margins are found to be an important process controlling the final topographic distribution of observed pahoehoe lobes.

modeling↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Boron Coordination in Multicomponent Glasses: Analytical Models and Machine Learning With Uncertainty

Borosilicate glasses are extensively used in a variety of applications from kitchenware to nuclear waste immobilization due to the strong network formed by the Si-O-B bond that makes it resistant to chemical corrosion and gives it a low thermal expansion. Boron, however, exists in both trigonal BO3 and tetrahedral BO4 bonds in glass systems, which impacts the chemical durability and thermal resistance of the glass, amongst other properties. Boron coordination (N4), or the ratio of the amount of BO4 to BO3 within a glass, may aid in predicting these properties but is difficult to derive without experimental data due to the complexity of impacts from varied glass compositions and processing factors. For this reason, compositional models have been developed to predict boron coordination, but the models typically include a limited number of glass components. To help fill this gap in the models, in this work, a diverse multicomponent glass dataset of 809 glasses is compiled from a literature search, and then a number of analytical and machine learning (ML) models are trained on the dataset. Previously developed modified Bernstein and modified Du Stebbins analytical models were fitted to update parameters with the new dataset. Then, partially Bayesian neural networks, Gaussian process regressor, and heteroskedastic deterministic neural networks were evaluated. The ML models examined all have different strategies to overcome the potential for overfitting as a result of a limited training dataset, and return results that account for model uncertainty, which can be valuable for understanding model reliability. For the first time, cooling rate is introduced as an input parameter for ML models, showing consistent improvements in performance and solidifying the importance of including parameters outside of composition alone for N4 prediction. The machine learning models examined here show promise in accurate predictions of boron coordination in borosilicate glasses, all achieving R2 values of 0.91.

boron coordination↗

Scalable computations for nonstationary Gaussian processes

Nonstationary Gaussian process models can capture complex spatially varying dependence structures in spatial datasets. However, the large number of observations in modern datasets makes fitting such models computationally intractable with conventional dense linear algebra. In addition, derivative-free or even first-order optimization methods can be very slow to converge when estimating many spatially varying parameters. In this paper, we present a computational framework which couples an algebraic block diagonal plus low-rank covariance matrix approximation with stochastic trace estimation to facilitate the efficient use of second-order solvers for maximum likelihood estimation of Gaussian process models with many parameters. We demonstrate the effectiveness of these methods by simultaneously fitting 192 parameters in the popular nonstationary model of Paciorek and Schervish using 107,600 sea surface temperature anomaly measurements.

97 MATHEMATICS AND COMPUTING↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Modeling Stochastic Variability in Multiband Time-series Data

In preparation for the era of time-domain astronomy with upcoming large-scale surveys, we propose a state-space representation of a multivariate damped random walk process as a tool to analyze irregularly-spaced multifilter light curves with heteroscedastic measurement errors. We adopt a computationally efficient and scalable Kalman filtering approach to evaluate the likelihood function, leading to maximum O(k 3 n) complexity, where k is the number of available bands and n is the number of unique observation times across the k bands. This is a significant computational advantage over a commonly used univariate Gaussian process that can stack up all multiband light curves in one vector with maximum O(k 3 n 3 ) complexity. Using such efficient likelihood computation, we provide both maximum likelihood estimates and Bayesian posterior samples of the model parameters. Three numerical illustrations are presented: (i) analyzing simulated five-band light curves for a comparison with independent single-band fits; (ii) analyzing five-band light curves of a quasar obtained from the Sloan Digital Sky Survey Stripe 82 to estimate short-term variability and timescale; (iii) analyzing gravitationally lensed g- and r-band light curves of Q0957+561 to infer the time delay. Two R packages, Rdrw and timedelay, are publicly available to fit the proposed models.

79 ASTRONOMY AND ASTROPHYSICS↗

Systematics in asteroseismic modelling: application of a correlated noise model for oscillation frequencies

ABSTRACT The detailed modelling of stellar oscillations is a powerful approach to characterizing stars. However, poor treatment of systematics in theoretical models leads to misinterpretations of stars. Here, we propose a more principled statistical treatment for the systematics to be applied to fitting individual mode frequencies with a typical stellar model grid. We introduce a correlated noise model based on a Gaussian process (GP) kernel to describe the systematics given that mode frequency systematics are expected to be highly correlated. We show that tuning the GP kernel can reproduce general features of frequency variations for changing model input physics and fundamental parameters. Fits with the correlated noise model better recover stellar parameters than traditional methods that either ignore the systematics or treat them as uncorrelated noise.

Li, Tanda (ORCID:0000000163962563)↗

Modelling populations of kilonovae

Abstract The 2017 detection of a kilonova coincident with gravitational-wave emission has identified neutron star mergers as the major source of the heaviest elements and dramatically constrained alternative theories of gravity. Observing a population of such sources has the potential to transform cosmology, nuclear physics, and astrophysics. However, with only one confident multi-messenger detection currently available, modelling the diversity of signals expected from such a population requires improved theoretical understanding. In particular, models that are quick to evaluate and are calibrated with more detailed multi-physics simulations are needed to design observational strategies for kilonovae detection and to obtain rapid-response interpretations of new observations. We use grey-opacity models to construct populations of kilonovae, spanning ejecta parameters predicted by numerical simulations. Our modelling focuses on wavelengths relevant for upcoming optical surveys, such as the Rubin Observatory Legacy Survey of Space and Time (LSST). In these simulations, we implement heating rates that are based on nuclear reaction network calculations. We create a Gaussian-process emulator for kilonova grey opacities, calibrated with detailed radiative transfer simulations. Using recent fits to numerical relativity simulations, we predict how the ejecta parameters from binary neutron star (BNS) mergers shape the population of kilonovae, accounting for the viewing-angle dependence. Our simulated population of BNS mergers produce peak i-band absolute magnitudes of −20 ≤ Mi ≤ −11. A comparison with detailed radiative transfer calculations indicates that further improvements are needed to accurately reproduce spectral shapes over the full light curve evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying experimental edge plasma evolution via multidimensional adaptive Gaussian process regression

The edge density and temperature of tokamak plasmas are strongly correlated with energy and particle confinement and their quantification is fundamental to understanding edge dynamics. These quantities exhibit behaviours ranging from sharp plasma gradients and fast transient phenomena (e.g. transitions between low and high confinement regimes) to nominal stationary phases. Analysis of experimental edge measurements therefore require robust fitting techniques to capture potentially stiff spatiotemporal evolution. Additionally, fusion plasma diagnostics inevitably involve measurement errors and data analysis requires a statistical framework to accurately quantify uncertainties. This paper outlines a generalized multidimensional adaptive Gaussian process routine capable of automatically handling noisy data and spatiotemporal correlations. We focus on the edge-pedestal region in order to underline advancements in quantifying time-dependent plasma profiles including transport barrier formation on the Alcator C-Mod tokamak.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Reconstructing the Universe: Testing the Mutual Consistency of the Pantheon and SDSS/eBOSS BAO Data Sets with Gaussian Processes

We test the mutual consistency between the baryon acoustic oscillation measurements from the eBOSS SDSS final release and the Pantheon supernova compilation in a model-independent fashion using Gaussian process regression. We also test their joint consistency with the ΛCDM model in a model-independent fashion. We also use Gaussian process regression to reconstruct the expansion history that is preferred by these two data sets. While this methodology finds no significant preference for model flexibility beyond ΛCDM, we are able to generate a number of reconstructed expansion histories that fit the data better than the best-fit ΛCDM model. These example expansion histories may point the way toward modifications to ΛCDM. We also constrain the parameters Ω{sub k} and H {sub 0} r {sub d} both with ΛCDM and with Gaussian process regression. We find that H {sub 0} r {sub d} = 10,030 ± 130 km s{sup −1} and Ω{sub k} = 0.05 ± 0.10 for ΛCDM and that H {sub 0} r {sub d} = 10,040 ± 140 km s{sup −1} and Ω{sub k} = 0.02 ± 0.20 for the Gaussian process case.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A machine learning approach for determining temperature-dependent bandgap of metal oxides utilizing Allen–Heine–Cardona theory and O’Donnell model parameterization

To evaluate the high temperature sensing properties of metal oxide and perovskite materials suitable for use in combustion environments, it is necessary to understand the temperature dependence of their bandgaps. Although such temperature-driven changes can be calculated via the Allen–Heine–Cardona (AHC) theory, which assesses electron–phonon coupling for the bandgap correction at given temperatures, this approach is computationally demanding. Another approach to predict bandgap temperature-dependence is the O’Donnell model, which uses analytical expressions with multiple fitting parameters that require bandgap information at 0 K. This work employs data-driven Gaussian process regression (GPR) to predict the parameters employed in the O’Donnell model from a set of physical features. We use a sample of 54 metal oxides for which density functional theory has been performed to calculate the bandgap at 0 K, and the AHC calculations have been carried out to determine the shift in the bandgap at non-zero temperatures. As the AHC calculations are impractical for high-throughput screening of materials, the developed GPR model attempts to alleviate this issue by predicting the O'Donnell parameters purely from physical features. To mitigate the reliability issues arising from the very small size of the dataset, we apply a Bayesian technique to improve the generalizability of the data-driven models as well as quantify the uncertainty associated with the predictions. The method captures well the overall trend of the O’Donnell parameters with respect to a reduced feature set obtained by transforming the available physical features. Quantifying the associated uncertainty helps us understand the reliability of the predictions of the O’Donnell parameters and, therefore, the bandgap as a function of temperature for any novel material.

36 MATERIALS SCIENCE↗

In Pursuit of the FIP Effect in Late-Type Stellar Coronae

Spectral line data for several coronally active stars, in addition to EUVE Deep Survey light curves, have been analysed under this program. Much difficulty has been encountered in the study that has resulted in fewer stars being analysed than had been hoped. The difficulties stemmed from the analysis of low X-ray spectra taken with the ASCA satellite that produced results that are strongly discrepant with respect to the EUVE results. There is no obvious explanation for this, though it appears that analysis of ASCA data systematically underestimate metal abundance in hot plasmas. Consequently, the final emphasis in our analyses has been on EUVE data. Observed line profiles have being fitted in order to measure their fluxes using IDL software specially developed under this and parallel efforts.The observed line profiles deviate from pure gaussian forms, but we have found the benefits of using additional functional forms in the fitting process to be of only very small value for the lines with highest S/N. The resulting line fluxes have being processed in terms of the coronal EM using new techniques. Resulting EM distribution models are being used to finalize metallicity and abundance estimates for the stars in the program. Special account of the influence of missing lines in the spectral models has been taken.

Drake, Jeremy↗

Fitting Matérn smoothness parameters using automatic differentiation

The Mat$\acute{e}$rn covariance function is ubiquitous in the application of Gaussian processes to spatial statistics and beyond. Perhaps the most important reason for this is that the smoothness parameter $\nu$ gives complete control over the mean-square differentiability of the process, which has significant implications for the behavior of estimated quantities such as interpolants and forecasts. Unfortunately, derivatives of the Mat$\acute{e}$rn covariance function with respect to $\nu$ require derivatives of the modified second-kind Bessel function $K$ $\nu$ with respect to $\nu$. While closed form expressions of these derivatives do exist, they are prohibitively difficult and expensive to compute. For this reason, many software packages require fixing $\nu$ as opposed to estimating it, and all existing software packages that attempt to offer the functionality of estimating $\nu$ use finite difference estimates for $\partial$ $\nu$ $K$ $\nu$ . In this work, we introduce a new implementation of $K$$\nu$ that has been designed to provide derivatives via automatic differentiation (AD), and whose resulting derivatives are significantly faster and more accurate than those computed using finite differences. Here, we provide comprehensive testing for both speed and accuracy and show that our AD solution can be used to build accurate Hessian matrices for second-order maximum likelihood estimation in settings where Hessians built with finite difference approximations completely fail.

97 MATHEMATICS AND COMPUTING↗

A Scalable Gaussian Process Approach to Shear Mapping with MuyGPs

Analysis of cosmic shear is an integral part of understanding structure growth across cosmic time, which in turn provides us with information about the nature of dark energy. Conventional methods generate shear maps from which we can infer the matter distribution in the universe. Current methods (e.g., Kaiser–Squires inversion) for generating these maps, however, are tricky to implement and can introduce bias. Recent alternatives construct a spatial process prior for the lensing potential, which allows for inference of the convergence and shear parameters given lensing shear measurements. Realizing these spatial processes, however, scales cubically in the number of observations—an unacceptable expense as near-term surveys expect billions of correlated measurements. Therefore, we present a linearly scaling shear map construction alternative using a scalable Gaussian process prior called MuyGPs. MuyGPs avoids cubic scaling by conditioning interpolation on only nearest neighbors and fits hyperparameters using batched leave-one-out cross-validation. This work is the first step toward a full, scalable mass mapping method. We work in a simplified regime where we validate our method by interpolating and analyzing maps given noisy point-estimate data from all three shear fields, taken from a suite of N -body ray-tracing simulations. We also show that we can perform these operations at the scale of billions of galaxies on high-performance computing platforms.

79 ASTRONOMY AND ASTROPHYSICS↗

Physics-informed machine learning modeling for predictive control using noisy data

Due to the occurrence of over-fitting at the learning phase, the modeling of chemical processes via artificial neural networks (ANN) by using corrupted data (i.e., noisy data) is an ongoing challenge. Therefore, this work investigates the effect of both Gaussian and non-Gaussian noise on the performance of process-structure based recurrent neural networks (RNN) models, which take the form of partially-connected RNN models in this work, that are used to approximate a class of multi-input-multi-outputs nonlinear systems. Furthermore, two different techniques, specifically Monte Carlo dropout and co-teaching, are utilized in the development of partially-connected RNN models. Here, these two techniques are employed to reduce the over-fitting in ANNs when noisy data is used in the training process and, hence, to improve the open-loop accuracy as well as the closed-loop performance under a Lyapunov-based model predictive controller (MPC). Aspen Plus Dynamics, a well-known high-fidelity process simulator, is used to simulate a large-scale chemical process application in order to demonstrate the anticipated improvements in both open-loop approximation and closed-loop controller performance in the presence of Gaussian and non-Gaussian noise in the data set using physics-informed RNNs.

97 MATHEMATICS AND COMPUTING↗

Physics-Based Machine Learning Methods for U-235 Forensics Signatures

Signatures of low-intensity U-235 sources have been recently studied by utilizing a variety of machine learning (ML) classifiers using features derived from gamma spectral measurements collected under structured campaigns. Several ML classifiers, such as ensemble of tress and classification trees, revealed misleadingly-optimistic training error due to over-fitting, and furthermore, their performance is not directly relatable to the physical properties due to their data-driven, opaque designs. We present a regression-based ML method that first estimates the inverse distance to the source and then utilizes a threshold to infer its presence, by representing the background as a source located at an infinite distance. For the inverse distance estimation, we study the ensemble of trees and Gaussian process regression methods, and a hyper parameter auto-tuning and selection method that employs five regression estimators. These methods avoid the over-fitting observed in several ML classifiers, while providing the classification error nearly comparable to them based on independent test data. Their error is directly related to estimates of the inverse physical distance to source, and the precision of error determines the seperability property that determines the false alarm and missed detection rates. The property of monotonic decrease of the source strength with increasing detector distance combined with Poisson distribution of measurements is utilized to analytically validate these methods by deriving the generalization equations of underlying regression methods.

Rao, Nageswara↗