Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Kernel density estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

Non-uniform active learning for Gaussian process models with applications to trajectory informed aerodynamic databases

The ability to non-uniformly weight the input space is desirable for many applications, and has been explored for space-filling approaches. Increased interests in linking models, such as in a digital twinning framework, increases the need for sampling emulators where they are most likely to be evaluated. In particular, here we apply non-uniform sampling methods for the construction of aerodynamic databases. This paper combines non-uniform weighting with active learning for Gaussian Processes (GPs) to develop a closed-form solution to a non-uniform active learning criterion. We accomplish this by utilizing a kernel density estimator as the weight function. We demonstrate the need and efficacy of this approach with an atmospheric entry example that accounts for both model uncertainty as well as the practical state space of the vehicle, as determined by forward modeling within the active learning loop.

42 ENGINEERING↗

Effects of wind field characteristics on pitch bearing reliability: a case study of 5 MW reference wind turbine at onshore and offshore sites; Auswirkungen von Windfeldeigenschaften auf die Zuverlässigkeit von Rotorblattlagern: eine Fallstudie über eine 5-MW-Referenz-Windkraftanlage an Onshore- und Offshore-Standorten

This paper presents a study on pitch bearing basic rating life affected by wind field characteristics at both onshore and offshore wind sites. The National Renewable Energy Laboratory 5 MW reference wind turbine is selected for the study. Wind field characteristics including reference hub height mean wind speed, wind speed distribution, wind shear, and vertical inflow are studied. A decoupled approach is employed where global analysis is performed first. Second, the load effects from the global analysis are applied on a reference pitch bearing designed based on best industrial practices. For the case study onshore site, it is found that the Kernel density estimation best fits the wind distribution, while the International Electrotechnical Commission proposed distribution appears to be not suitable. Moreover, it is shown that the seed number has high effect on the bearing life in turbulence wind and the wind speeds around rated have the highest contribution in both bearing fatigue damage and extreme load failure. The results contribute to better understanding of the wind field characteristics on the pitch bearing life.

17 WIND ENERGY↗

Benchmarking FFTF LOFWOS Test# 13 using SAM code: Baseline model development and uncertainty quantification

The development and deployment of advanced reactors, such as the sodium-cooled fast reactor (SFR), relies on sophisticated modeling tools to ensure the safety of the design under various transients. The predictive capability of these advanced modeling tools requires validation to garner trust in supporting the licensing of the advanced reactors. For this reason, the International Atomic Energy Agency (IAEA) initiated a coordinated research project (CRP) in 2018 for the analysis of the Fast Flux Test Facility (FFTF) Loss of Flow Without Scram (LOFWOS) Test #13.In this study, we present and discuss the benchmarking efforts of the modern system code SAM on the FFTF LOFWOS Test #13. Further, the SAM baseline model was developed according to the benchmark specification, which included a detailed core model with reactivity feedback. Generally, good agreement was observed between the baseline results and benchmark measurements; however, discrepancies persisted, particularly in predicted fuel assembly coolant outlet temperatures. Utilizing the baseline model, uncertainty quantification (UQ) and sensitivity analysis (SA) were conducted with the assistance of various statistical learning and machine learning methods, including kernel density estimation, Gaussian processes, and Sobol indices. Following the baseline model prediction and UQ and SA results, we discuss the reasons for the simulation discrepancies and propose further improvements to the model. This benchmarking effort adheres to the best-estimate plus uncertainty approach and can serve as a valuable example for supporting risk-informed licensing of advanced reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

A meshless stochastic method for Poisson–Nernst–Planck equations

A plethora of biological, physical, and chemical phenomena involve transport of charged particles (ions). Its continuum-scale description relies on the Poisson–Nernst–Planck (PNP) system, which encapsulates the conservation of mass and charge. The numerical solution of these coupled partial differential equations is challenging and suffers from both the curse of dimensionality and difficulty in efficiently parallelizing. We present a novel particle-based framework to solve the full PNP system by simulating a drift–diffusion process with time- and space-varying drift. We leverage Green’s functions, kernel-independent fast multipole methods, and kernel density estimation to solve the PNP system in a meshless manner, capable of handling discontinuous initial states. The method is embarrassingly parallel, and the computational cost scales linearly with the number of particles and dimension. We use a series of numerical experiments to demonstrate both the method’s convergence with respect to the number of particles and computational cost vis-à-vis a traditional partial differential equation solver.

Chemistry↗

Reionization effective likelihood from Planck 2018 data

We release relike (reionization effective likelihood), a fast and accurate effective likelihood code based on the latest Planck 2018 data that allows one to constrain any model for reionization between 6 < z < 30 using five constraints from the CMB reionization principal components (PC). We tested the code on two example models which showed excellent agreement with sampling the exact Planck likelihoods using either a simple Gaussian PC likelihood or its full kernel density estimate. This code enables a fast and consistent means for combining Planck constraints with other reionization data sets, such as kinetic Sunyaev-Zeldovich effects, line-intensity mapping, luminosity function, star formation history, quasar spectra, etc., where the redshift dependence of the ionization history is important. Since the PC technique tests any reionization history in the given range, we also derive model-independent constraints for the total Thomson optical depth τ PC = $0.0619$$^{+0.0056}_{–0.0068}$ and its 15 ≤ z ≤ 30 high redshift component τ PC (15,30) < 0.020 (95% C.L.). Furthermore, the upper limits on the high-redshift optical depth is a factor of ~3 larger than those reported in the Planck 2018 cosmological parameter paper using the FlexKnot method and we validate our results with a direct analysis of a two-step model which permits this small high-z component.

79 ASTRONOMY AND ASTROPHYSICS↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Next-Cycle Optimal Fuel Control for Cycle-to-Cycle Variability Reduction in EGR-Diluted Combustion

In this simulation study, cycle-to-cycle fuel control was used to reduce CCV by injecting additional fuel in operating conditions with sporadic misfires and partial burns. An optimal control policy was proposed that utilizes 1) a physics-based model that tracks in-cylinder gas composition and 2) a one-step-ahead prediction of the combustion efficiency based on a kernel density estimator. The optimal solution, however, presents a tradeoff between the reduction in combustion CCV and the increase in fuel injection quantity required to stabilize the charge. Such a tradeoff can be ad- just by a single parameter embedded in the cost function.

Maldonado, BryanP. [Oak Ridge National Lab. (ORNL)↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

Polynomial Chaos Surrogate Construction for Random Fields with Parametric Uncertainty

Engineering and applied science rely on computational experiments to rigorously study physical systems. The mathematical models used to probe these systems are highly complex, and sampling-intensive studies often require prohibitively many simulations for acceptable accuracy. Surrogate models provide a means of circumventing the high computational expense of sampling such complex models. In particular, polynomial chaos expansions (PCEs) have been successfully used for uncertainty quantification studies of deterministic models where the dominant source of uncertainty is parametric. We discuss an extension to conventional PCE surrogate modeling to enable surrogate construction for stochastic computational models that have intrinsic noise in addition to parametric uncertainty. We develop a PCE surrogate on a joint space of intrinsic and parametric uncertainty, enabled by Rosenblatt transformations, which are evaluated via kernel density estimation of the associated conditional cumulative distributions. Furthermore, we extend the construction to random field data via the Karhunen–Loève expansion. We then take advantage of closed-form solutions for computing PCE Sobol indices to perform a global sensitivity analysis of the model which quantifies the intrinsic noise contribution to the overall model output variance. Additionally, the resulting joint PCE is generative in the sense that it allows generating random realizations at any input parameter setting that are statistically approximately equivalent to realizations from the underlying stochastic model. The method is demonstrated on a chemical catalysis example model and a synthetic example controlled by a parameter that enables a switch from unimodal to bimodal response distributions.

97 MATHEMATICS AND COMPUTING↗

PV-Finder: ML Based Algorithm for Primary Vertex Identification

he CMS detector at the High-Luminosity Large Hadron Collider (HL-LHC) will operate in challenging conditions with expected pile-up of up to 200 collisions per bunch crossing, necessitating the development of a more resilient primary vertex (PV) reconstruction method to ensure the integrity of data analysis and the efficiency of the CMS triggering system. This contribution describes preliminary studies on a new ML based PV-Finder method for PV identification. The method is based on a model trained using Kernel Density Estimations (KDEs) derived from the positions of reconstructed tracks at the beamline, incorporating uncertainties from track parameters. It also utilizes target histograms, modeled as Gaussian distributions centered on the actual ground truth values of specific primary vertices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Non-Parametric Statistical Analysis of Current Waveforms through Power System Sensors

The protection, control, and monitoring of the power grid is not possible without accurate measurement devices. As the percentage of renewable energy sources penetrating the existing grid infrastructure increases, so do uncertainties surrounding their effects on the everyday operation of the power system. Many of these devices are sources of high-frequency transients. These transients may be useful for identifying certain events or behaviors otherwise not seen in traditional analysis techniques. Therefore, the ability of sensors to accurately capture these phenomena is paramount. In this work, two commercial-grade power system distribution sensors are investigated in terms of their ability to replicate high-frequency phenomena by studying their responses to three events: a current inrush, a microgrid “close-in”, and a fault on the terminals of a wind turbine. Kernel density estimation is used to derive the non-parametric probability density functions of these error distributions and their adequateness is quantified utilizing the commonly used root mean square error (RMSE) metric. It is demonstrated that both sensors exhibit characteristics in the high harmonic range that go against the assumption that measurement error is normally distributed.

47 OTHER INSTRUMENTATION↗

Learning User Preferences for Sets of Objects

Most work on preference learning has focused on pairwise preferences or rankings over individual items. In this paper, we present a method for learning preferences over sets of items. Our learning method takes as input a collection of positive examples--that is, one or more sets that have been identified by a user as desirable. Kernel density estimation is used to estimate the value function for individual items, and the desired set diversity is estimated from the average set diversity observed in the collection. Since this is a new learning problem, we introduce a new evaluation methodology and evaluate the learning method on two data collections: synthetic blocks-world data and a new real-world music data collection that we have gathered.

preferences↗

Pattern Recognition for a Flight Dynamics Monte Carlo Simulation

The design, analysis, and verification and validation of a spacecraft relies heavily on Monte Carlo simulations. Modern computational techniques are able to generate large amounts of Monte Carlo data but flight dynamics engineers lack the time and resources to analyze it all. The growing amounts of data combined with the diminished available time of engineers motivates the need to automate the analysis process. Pattern recognition algorithms are an innovative way of analyzing flight dynamics data efficiently. They can search large data sets for specific patterns and highlight critical variables so analysts can focus their analysis efforts. This work combines a few tractable pattern recognition algorithms with basic flight dynamics concepts to build a practical analysis tool for Monte Carlo simulations. Current results show that this tool can quickly and automatically identify individual design parameters, and most importantly, specific combinations of parameters that should be avoided in order to prevent specific system failures. The current version uses a kernel density estimation algorithm and a sequential feature selection algorithm combined with a k-nearest neighbor classifier to find and rank important design parameters. This provides an increased level of confidence in the analysis and saves a significant amount of time.

Restrepo, Carolina↗

Using Spatial Density to Characterize Volcanic Fields on Mars

We introduce a new tool to planetary geology for quantifying the spatial arrangement of vent fields and volcanic provinces using non parametric kernel density estimation. Unlike parametricmethods where spatial density, and thus the spatial arrangement of volcanic vents, is simplified to fit a standard statistical distribution, non parametric methods offer more objective and data driven techniques to characterize volcanic vent fields. This method is applied to Syria Planum volcanic vent catalog data as well as catalog data for a vent field south of Pavonis Mons. The spatial densities are compared to terrestrial volcanic fields.

Richardson, J. A.↗

Architecture and Dynamics of Kepler's Multi-Transiting Planet Systems.III. Comprehensive Investigation Using All Four Years of Kepler Mission Data

We study the orbital architectures of planetary systems orbiting within ∼1 AU of their stars by analyzing the ensemble of Kepler systems having two or more planet candidates. We use data from the entire Kepler mission, and in many cases we apply improved analysis techniques (e.g., replacing histograms by top-hat Kernel Density Estimators that avoid the loss of information resulting from choosing a particular phase for the bin boundaries) to extend and enhance the studies of Lissauer et al. (2011, ApJS 197, 8) and Fabrycky et al. (2014, ApJ 790, 146).These data show ~ 1700 transiting planet candidates in > 600 multiple-planet systems, far more than were available for our previous two studies. The increased numbers and better information about planetary radii and the properties of stellar hosts made possible by Gaia DR2 allow more statistically-robust analyses of the entire ensemble of Kepler multis as well as independent analyses of subsets of the population. We are thus able to contrast the dynamical configurations of small and large planets, short-period and longer-period planets, and planets orbiting various types of host stars. We reinforce our previous findings that most pairs of planets within the same system are neither in nor near low-order mean motion resonances and that there is a substantial excess of planets having period ratios slightly larger than those of first-order mean-motion resonances. However, neglecting three systems whose planets are locked in 3- body resonances and summing over all first-order mean motion resonances, the deficit of planet pairs with period ratios just narrow of resonance is as large as the excess of planets wide of resonance (within statistical uncertainties), suggesting that overall there is no overall excess of planet pairs in the vicinity of resonance. Other aspects of our study, including estimates of the typical relative inclinations of planetary orbits and their variations as functions of orbital period, planet sizes and stellar properties, are in progress, with results expected to be available for presentation by the time of the conference.

Lissauer, Jack↗