Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian processes regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Visualization and Quantification of Wind Induced Variability in Hydrogen Clouds Following Releases of Liquid Hydrogen: Preprint

Well characterized experimental data for consequence model validation is important in progressing the use of liquid hydrogen as an energy carrier. In 2019, the Health and Safety Executive (HSE) undertook a series of liquid hydrogen dispersion and combustion experiments as a part of the Pre-normative Research into the Safe Use of Liquid Hydrogen (PRESLHY) project. In partnership between the National Renewable Energy Laboratory (NREL) and HSE, time and spatially varying hydrogen concentration measurements were made in 25 dispersion experiments and 23 congested ignition experiments associated with PRESLHY WP3 and WP5, respectively. These measurements were undertaken using the hydrogen wide area monitoring system developed by NREL. During the 23 congested ignition experiments, high variability was observed in the measured explosion severity during experiments with similar initial conditions. This led to the conclusion that wind, including localized gusts, had a large influence on the dispersion of the hydrogen, and therefore the quantity of hydrogen that was present in the congested region of the explosions. Using the hydrogen concentration measurements taken immediately prior to ignition, the hydrogen clouds were visualized in an attempt to rationalize the variability in overpressure between the tests. Gaussian process regression was applied to quantify the variability of the measured hydrogen concentrations. This analysis could also be used to guide modifications in experimental designs for future research on hydrogen combustion behavior.

HSR&D↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Estimators and Fusers for Fiber Delay Estimation Using Environmental Measurements

The properties of deployed network fiber are affected by environmental factors due to their exposure to the elements. Particularly for quantum networks, the resultant delay variations may have significant impacts due to the extreme sensitivity of synchronization, coincidence counting, and other critical operations. In this paper, the delays of 15 km aerial-inground fiber connections are measured, and effects due to temperature, humidity and wind speed are analyzed over multiple periods spanning four seasons of a year. Machine learning methods are first utilized to reveal surprisingly pronounced effects of humidity on the delay, in addition to the expected temperature and its seasonal variations. Estimator and fusion methods are developed to estimate the delay using temperature, humidity and wind speed measurements, by utilizing smooth Gaussian Process Regression (GPR) and nonsmooth Ensemble of Trees (EOT) methods. Measurements from winter and summer periods are temporally fused using twelve different methods, and eight methods provide estimates for the delay throughout the year with median test errors under 1.28%. The results reveal distinct temperature-humidity trends across the seasons, and the ability of estimator and temporal fusion methods to exploit them for estimating the delay. These results constitute a case study of machine learning analytical results, wherein generalization equations explain the performance of various estimator and fuser methods.

Rao, Nageswara [ORNL] (ORCID:0000000234085941)↗

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING↗

A Gaussian process guide for signal regression in magnetic fusion

Extracting reliable information from diagnostic data in tokamaks is critical for understanding, analyzing, and controlling the behavior of fusion plasmas and validating models describing that behavior. Recent interest within the fusion community has focused on the use of principled statistical methods, such as Gaussian process regression (GPR), to attempt to develop sharper, more reliable, and more rigorous tools for examining the complex observed behavior in these systems. While GPR is an enormously powerful tool, there is also the danger of drawing fragile, or inconsistent conclusions from naive GPR fits that are not driven by principled treatments. Here we review the fundamental concepts underlying GPR in a way that may be useful for broad-ranging applications in fusion science. We also revisit how GPR is developed for profile fitting in tokamaks. We examine various extensions and targeted modifications applicable to experimental observations in the edge of the DIII-D tokamak. Finally, we discuss best practices for applying GPR to fusion data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Online Learning of Effective Turbine Wind Speed in Wind Farms

To develop better wind farm controllers that can meet more complex objectives, methods of modeling the wind turbine wakes at low computational expense are needed. Gaussian process (GP) regression offers a computationally inexpensive framework for learning complex functions from noisy measurements with very few datapoints. In this work, an online learning approach is presented to learn the rotor-averaged wind velocity at downstream wind turbines with GPs, using the available datastream of wind field measurements and wind turbine control set-points. This framework can readily be integrated into model-based controls methods because the model a) is updated online at low computational expense, b) assumes a mathematically favorable Gaussian form, and c) explicitly quantifies the stochastic nature of the wake field so that the trade-off between exploration and exploitation, and the uncertainty in the prediction, can be utilized. We show that a GP-learned model can match true values with errors within 0.5% on average, with as few as 5 training data points.

Gaussian process↗

O’Hare Airport Short-Term Ground Transportation Modal Demand Forecast Using Gaussian Processes

Here, the principal objective of this study is to analyze the spatial and temporal variation of ground transportation airport demand and provide demand forecast to inform planning capability and explore alternatives for investments to accommodate airport growth. Because of its good adaptability and strong generalization ability for dealing with high-dimensional input, small-sample, and nonlinear spatial data, Gaussian process (GP) regression is used to provide forecast estimates using data from transportation network company (TNC) trips and urban rail passengers at Chicago's O'Hare International Airport. TNC airport trips differ significantly, with three times more distance, more than twice the travel time, and half of the share requests compared with nonairport trips. This highlights the need for separate demand models. Hourly analysis of the rail service indicates that this is likely heavily used by airport workers, whereas TNC services focus on travelers because of variations in the peak demand hours. Heteroscedastic GP regression is implemented because of differences in trip variance between night and day hours. Estimates are given for weekdays and weekend trips, and the 95% confidence intervals are calculated. The introduction of flight schedule information into the models shows marginal improvements in their performance. However, fitting a GP regression becomes computationally expensive with increased sample size and the introduction of spatial components. Transportation planners and policymakers can use the results and methods implemented in this study to optimize transportation assets and provide long-range simulations of the current and future conditions in the area.

42 ENGINEERING↗

Learning-Based Demand Response in Grid Interactive Buildings via Gaussian Processes

This paper presents a predictive controller for a grid-interactive multi-zone building where the temperature dynamics are learned via Gaussian Process (GP) regression. We investigate the development of a learning-based predictive control with two main objectives: (i) continuously learn the temperature dynamics of the building based on data; and, (ii) use the learned dynamics to solve a multi-objective predictive control problem to guarantee occupants' comfort and energy efficiency during normal conditions and demand response events. We leverage the probabilistic non-parametric properties of GPs to estimate the (unknown) non-linear temperature dynamics of the building and to incorporate the uncertainty of those predictions in a multi-objective optimization problem. The GP-based predictive control is solved via a zero-order primal-dual projected-gradient algorithm. We evaluate numerically the performance of the proposed controller using a five-zone commercial building.

demand response↗

Detection Limits of Low-mass, Long-period Exoplanets Using Gaussian Processes Applied to HARPS-N Solar Radial Velocities

Radial velocity (RV) searches for Earth-mass exoplanets in the habitable zone around Sun-like stars are limited by the effects of stellar variability on the host star. In particular, suppression of convective blueshift and brightness inhomogeneities due to photospheric faculae/plage and starspots are the dominant contribution to the variability of such stellar RVs. Gaussian process (GP) regression is a powerful tool for statistically modeling these quasi-periodic variations. We investigate the limits of this technique using 800 days of RVs from the solar telescope on the High Accuracy Radial velocity Planet Searcher for the Northern hemisphere (HARPS-N) spectrograph. These data provide a well-sampled time series of stellar RV variations. Into this data set, we inject Keplerian signals with periods between 100 and 500 days and amplitudes between 0.6 and 2.4 m s{sup −1}. We use GP regression to fit the resulting RVs and determine the statistical significance of recovered periods and amplitudes. We then generate synthetic RVs with the same covariance properties as the solar data to determine a lower bound on the observational baseline necessary to detect low-mass planets in Venus-like orbits around a Sun-like star. Our simulations show that discovering planets with a larger mass (∼0.5 m s{sup −1}) using current-generation spectrographs and GP regression will require more than 12 yr of densely sampled RV observations. Furthermore, even with a perfect model of stellar variability, discovering a true exo-Venus (∼0.1 m s{sup −1}) with current instruments would take over 15 yr. Therefore, next-generation spectrographs and better models of stellar variability are required for detection of such planets.

47 OTHER INSTRUMENTATION↗

K-means-driven Gaussian Process data collection for angle-resolved photoemission spectroscopy

Abstract We propose the combination of k-means clustering with Gaussian Process (GP) regression in the analysis and exploration of 4D angle-resolved photoemission spectroscopy (ARPES) data. Using cluster labels as the driving metric on which the GP is trained, this method allows us to reconstruct the experimental phase diagram from as low as 12% of the original dataset size. In addition to the phase diagram, the GP is able to reconstruct spectra in energy-momentum space from this minimal set of data points. These findings suggest that this methodology can be used to improve the efficiency of ARPES data collection strategies for unknown samples. The practical feasibility of implementing this technology at a synchrotron beamline and the overall efficiency implications of this method are discussed with a view on enabling the collection of more samples or rapid identification of regions of interest.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Predicting the Activity and Selectivity of Bimetallic Metal Catalysts for Ethanol Reforming using Machine Learning

Machine learning is ideally suited for the pattern detection in large uniform datasets, but consistent experimental datasets on catalyst studies are often small. Here we demonstrate how a combination of machine learning and first-principles calculations can be used to extract knowledge from a relatively small set of experimental data. The approach is based on combining a complex machine-learning model trained on an extensive computational library of transition-state energies with simple linear regression models of experimental catalytic activities and selectivities from the literature. Using the combined model, we identify the key C–C bond scission reactions involved in ethanol reforming and perform a computational screening for ethanol reforming on monolayer bimetallic catalysts with architectures TM-Pt-Pt(111) and Pt-TM-Pt(111) (TM = 3d transition metals). The model also predicts four promising catalyst compositions for future experimental studies. In conclusion, the approach is not limited to ethanol reforming but is of general use for the interpretation of experimental observations as well as for the computational discovery of novel catalytic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A surrogate-model-based approach for estimating the first and second-order moments of offshore wind power

Power curve, the functional relationship that governs the process of converting a set of weather variables experienced by a wind turbine into electric power, is widely used in the wind industry to estimate power output for planning and operational purposes. Existing methods for power curve estimation have three main limitations: (i) they mostly rely on wind speed as the sole input, thus ignoring the secondary, yet possibly significant effects of other environmental factors, (ii) they largely overlook the complex marine environment in which offshore turbines operate, potentially compromising their value in offshore wind energy applications, and (ii) they solely focus on the first-order properties of wind power, with little (or null) information about the variation around the mean behavior, which is important for ensuring reliable grid integration, asset health monitoring, and energy storage, among others. In light of that, this study investigates the impact of several wind-and wave-related factors on offshore wind power variability, with the ultimate goal of accurately predicting its first two moments. Further, our approach couples OpenFAST—a multi-physics wind turbine simulator—with Gaussian Process (GP) regression to reveal the underlying relationships governing offshore weather-to-power conversion. We first find that a multi-input power curve which captures the combined impact of wind speed, direction, and air density, can provide double-digit improvements, in terms of prediction accuracy, relative to univariate methods which rely on wind speed as the sole explanatory variable (e.g. the standard method of bins). Wave-related variables are found not important for predicting the average power output, but interestingly, appear to be extremely relevant in describing the fluctuation of the offshore power around its mean. Tested on real-world data collected at the New York/New Jersey bight, our proposed multi-input models demonstrate a high explanatory power in predicting the first two moments of offshore wind generation, testifying their potential value to the offshore wind industry.

17 WIND ENERGY↗

CAMERA: A method for cost-aware, adaptive, multifidelity, efficient reliability analysis

Estimating probability of failure in aerospace systems is a critical requirement for flight certification and qualification. Failure probability estimation involves resolving tails of probability distributions, and Monte Carlo sampling methods are intractable when expensive high-fidelity simulations have to be queried. Here, we propose a method to use models of multiple fidelities that trade accuracy for computational efficiency. Specifically, we propose the use of multifidelity Gaussian process models to efficiently fuse models at multiple fidelity, thereby offering a cheap surrogate model that emulates the original model at all fidelities. Furthermore, we propose a novel sequential acquisition function based experiment design framework that can automatically select samples from appropriate fidelity models to make predictions about quantities of interest at the highest fidelity. We use our proposed approach in an importance sampling setting and demonstrate our method on the failure level set and probability estimation on synthetic test functions and two real-world applications, namely, the reliability analysis of a gas turbine engine blade using a finite element method and a transonic aerodynamic wing test case using Reynolds-averaged Navier-Stokes equations. We show that our method predicts the failure boundary and probability more accurately and at a fraction of the computational cost compared with using just a single expensive high-fidelity model. Finally, we show that our sequential approach is guaranteed to asymptotically converge to the true failure boundary with high probability.

97 MATHEMATICS AND COMPUTING↗

Predicting boron coordination in multicomponent borate and borosilicate glasses using analytical models and machine learning

Accurate prediction of boron coordination in multicomponent glasses is critical in glass science and technology as it strongly affects the properties of borate and borosilicate glasses. We have collected a dataset containing 657 glasses from literature with boron coordination values and developed models using analytical functions based on the well accepted Dell, Xiao and Bray model. Good prediction of boron coordination with a R 2 value higher than 0.8 was obtained. The large variation of boron coordination from experiments, originated from sample preparations and characterizations, led to difficulties in obtaining models with better prediction performance. Various machine learning (ML) algorithms were evaluated and slightly better prediction performance was observed; however, interpretation of the ML models is less straight forward. In conclusion, this study developed various models capable of providing quantitative boron coordination predictions, providing insights into its structural roles in multi-component glasses, and suggesting fruitful areas for future research.

36 MATERIALS SCIENCE↗

Transient uncertainty quantification and Global Sensitivity Analysis of the open-source Molten Chloride Reactor Experiment (MCRE) using GP-PCA surrogate models

Uncertainties in the thermophysical properties of molten salts impact both the steady-state and transient behavior of Molten Salt Reactors (MSRs). In this work, we aim to quantify the influence of such uncertainties on the transient operation of the Molten Chloride Reactor Experiment (MCRE), utilizing the open-source specifications provided for this reactor. Seven representative transient scenarios are considered. For each scenario, we evaluate the impact of thermophysical property uncertainties on four key multiphysics model output variables of interest (VoIs): maximum power density, maximum fuel temperature, maximum reflector temperature, and average fuel velocity magnitude. In addition, we perform a Global Sensitivity Analysis (GSA) by computing Sobol’ indices for the uncertain input parameters to determine their contribution to the variability of each VoI. Conducting GSA is computationally intensive due to the large number of required evaluations of the high-fidelity multiphysics model. To mitigate this cost, we develop a surrogate modeling framework that combines Gaussian Process (GP) regression with Principal Component Analysis (PCA), enabling efficient sample generation for the GSA. Our results show that for energy-related VoIs, thermal conductivity is the dominant contributor to uncertainty. In contrast, for flow-related VoIs, density and dynamic viscosity are the primary sources of uncertainty. The specific heat of the fuel salt was found to play a secondary role in the transient analyses.

42 - ENGINEERING↗