Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel density estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Density Estimation with Mercer Kernels

We present a new method for density estimation based on Mercer kernels. The density estimate can be understood as the density induced on a data manifold by a mixture of Gaussians fit in a feature space. As is usual, the feature space and data manifold are defined with any suitable positive-definite kernel function. We modify the standard EM algorithm for mixtures of Gaussians to infer the parameters of the density. One benefit of the approach is it's conceptual simplicity, and uniform applicability over many different types of data. Preliminary results are presented for a number of simple problems.

Macready, William G.↗

Efficient screening of rare large pit anomalies on polished surfaces using a minimalist sampling scheme

Lawrence Livermore National Laboratory (LLNL) has made significant strides in generating clean energy through its inertial confinement fusion (ICF) experiments. These experiments rely on high-density carbon (HDC) coated shells to encapsulate the fusion fuel. The success of these experiments is heavily dependent on the surface quality of these shells, as even minor imperfections, such as deep pits, can negatively impact fusion yield. Ensuring the required smoothness involves an extensive surface-finishing process that spans approximately 20 stages, making it both time-intensive and resource-demanding. A critical challenge in this process is the need for high-resolution scans to detect rare deep pits, which can be costly and impractical if performed on every shell. This highlights the necessity of developing more efficient scanning methods to optimize time and cost without compromising accuracy. To address these challenges, we introduce a novel approach that employs the multivariate Dvoretzky–Kiefer–Wolfowitz (DKW) inequality to provide a probabilistic upper bound on the error in estimating pit distribution characteristics via a Kernel Density Estimator (KDE). This error bound enables efficient and reliable estimation of pit distribution characteristics at a specified statistical confidence level using a minimal number of surface scans. The integrated DKW-KDE approach was validated through surface-finishing experiments across two batches of HDC-coated shells, demonstrating consistent and robust performance across multiple stages of the surface-finishing experiments. The validation studies suggest that the integrated DKW-KDE approach achieves comparable accuracy in estimating the risk of deleterious large pits with six scans, thus conserving time and resources. Further evaluations show that performance remains consistent across batches and over multiple polishing stages. In conclusion, based on these findings, one can leverage the minimal-scan insights to strategically improve the bottleneck inspection process, thus enhancing the productivity and quality of shell polishing and similar challenging manufacturing processes.

Inertial confinement fusion↗

A Multi-Fidelity Gaussian Process Regression Method for Probabilistic Wind Farm Power Curve Estimation

Accurate estimation of the power curve for wind turbines or wind farms is crucial to ensure their efficient operation and management. However, conventional methods for power curve estimation rely either on expensive and infrequent measurements or on low-quality numerical simulations. Moreover, the majority of previous studies on power curve estimation for wind turbines or wind farms focused on deterministic estimation, which provides a point estimate of the relationship between wind speed and power generation. Nevertheless, the deterministic approach fails to consider the inherent uncertainty associated with wind energy production resulting from varying turbine characteristics. This can lead to inaccurate power generation estimation and suboptimal decisions regarding energy management. In this paper, a kernel density estimation (KDE) based Multi-Fidelity Gaussian Process Regression (MFGPR) model is proposed to fuse theoretical power curve data and the ground true measurements to create a mapping of wind speed and wind power. By conducting a case study on an actual wind farm in China, the efficacy of the proposed MFGPR model was demonstrated in characterizing the variability of wind power. The probabilistic MFGPR model was also able to generate confidence intervals that encompassed the measured power, thereby improving the accuracy and confidence in wind power estimation or wind resource assessment. Overall, the proposed MFGPR model offers a reliable approach to integrate high-fidelity ground measurements and theoretical power curve data, resulting in precise wind resource assessment and power estimation.

Gaussian process regression↗

A functional global sensitivity measure and efficient reliability sensitivity analysis with respect to statistical parameters

Sensitivity analysis and reliability assessment are two important aspects of structural and system safety. Epistemic uncertainty with respect to probabilistic model of input parameters due to lack of knowledge is present in many scarce-data applications and complicates the characterization of uncertainty in model response. In this article, we present two importance measures to evaluate the impact of distribution parameters on the probability distribution function (PDF) of the output and the failure probability. The epistemic uncertainty associated with the distribution parameters is modeled as random variables. Additionally, a modified extended polynomial chaos expansion (MEPCE) approach is introduced in which aleatory and epistemic random variables are modeled and propagated simultaneously while allowing the separate assessment for any single epistemic variable. A MEPCE-based kernel density estimation (KDE) construction provides a composite map from each epistemic variable to the response PDF. The functional global sensitivity index of the PDF with respect to the distribution parameters is thus derived, as a function of output, which is both more informative and more efficient than standard scalar sensitivity measures. Reliability sensitivity indices can be readily evaluated by integrating the global sensitivity index function over the failure zone. Three illustrative examples are used to demonstrate the proposed methodology.

42 ENGINEERING↗

Improved Data Interpretation through Identification of Time Series Periodicity Changes

Analysis and interpretation of time series data is easiest when the data values occur at uniform intervals in time, but actual data may have differing data sampling frequencies, such as monthly and daily readings. Applying data analysis techniques, such as smoothing, to such a data set may not give a representative result between time segments. The ability to automatically distinguish time segments of differing data frequency would provide a means for applying data analysis independently to each segment, though a suitable blending at segment boundaries would be required. A method for detecting frequency changes was developed and applied to Gaussian and median smoothing of hydraulic head data from groundwater wells at the U.S. Department of Energy Hanford Site in southeastern Washington state. The process identifies time segments of high-frequency (daily) or low-frequency (greater than daily) data using adjusted-bandwidth Gaussian kernel density estimation and a threshold value, which are further refined to address small blocks of low-frequency data within larger blocks of high-frequency data. User-selectable levels of smoothing are then applied independently to the time segments prior to combining the segment results for a single smoothed data set. This time segment identification approach provides effective low- and high-frequency data separation, which provides a method to apply data analysis independently to each time segment.

97 MATHEMATICS AND COMPUTING↗

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

NASA Langley's Approach to the Sandia's Structural Dynamics Challenge Problem

The objective of this challenge is to develop a data-based probabilistic model of uncertainty to predict the behavior of subsystems (payloads) by themselves and while coupled to a primary (target) system. Although this type of analysis is routinely performed and representative of issues faced in real-world system design and integration, there are still several key technical challenges that must be addressed when analyzing uncertain interconnected systems. For example, one key technical challenge is related to the fact that there is limited data on target configurations. Moreover, it is typical to have multiple data sets from experiments conducted at the subsystem level, but often samples sizes are not sufficient to compute high confidence statistics. In this challenge problem additional constraints are placed as ground rules for the participants. One such rule is that mathematical models of the subsystem are limited to linear approximations of the nonlinear physics of the problem at hand. Also, participants are constrained to use these models and the multiple data sets to make predictions about the target system response under completely different input conditions. Our approach involved initially the screening of several different methods. Three of the ones considered are presented herein. The first one is based on the transformation of the modal data to an orthogonal space where the mean and covariance of the data are matched by the model. The other two approaches worked solutions in physical space where the uncertain parameter set is made of masses, stiffnesses and damping coefficients; one matches confidence intervals of low order moments of the statistics via optimization while the second one uses a Kernel density estimation approach. The paper will touch on all the approaches, lessons learned, validation 1 metrics and their comparison, data quantity restriction, and assumptions/limitations of each approach. Keywords: Probabilistic modeling, model validation, uncertainty quantification, kernel density

Horta, Lucas G.↗

GalaxyFlow: upsampling hydrodynamical simulations for realistic mock stellar catalogues

ABSTRACT Cosmological N-body simulations of galaxies operate at the level of ‘star particles’ with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogues requires ‘upsampling’ the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase-space density and sample from it. Secondly, we improve on existing upsamplers based on adaptive kernel density estimation (KDE), using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighbourhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalogue. Furthermore, we introduce a novel multimodel classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow more accurately estimates the density of the underlying star particles than methods based on KDE, at the cost of being more computationally intensive.

Lim, Sung Hak (ORCID:0000000330981092)↗

Experimental and Computational Evaluation of Lipidomic In-Source Fragmentation as a Result of Postionization with Matrix-Assisted Laser Desorption/Ionization

Matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI) can provide spatially resolved molecular information about a sample. Recently, a postionization approach (MALDI-2) has been commercially integrated with MALDI-MSI, allowing for bettered sensitivity and consequent improved spatial resolution. While advantages of MALDI-2 have previously been established, we demonstrate here statistically increased in-source fragmentation (ISF) results from postionization with a commercial instrument. Via lipid standard analyses, known MALDI ISF pathways (e.g., loss of trimethylamine) were statistically increased in MALDI-2 compared to MALDI-1 (65–172% increase in fragmentation). Gas phase molecular modeling with density functional theory estimated that the most-weighted virtual orbitals to excite within lipids involve ester and phosphate bonds. Protonated lipid excitation energies are furthermore red-shifted compared to those of other adduct types [e.g., 254 nm for protonated PC(16:0/18:1)] and approach the MALDI-2 laser energy (266 nm). Analysis of rat brain homogenate detected statistically more positive-ion mode peaks with MALDI-2 (1090) than that with MALDI-1 (719), where Kernel density estimations showed that the majority of this enhancement occurs with low m/z ions (i.e., m/z 75–500). Taken together with the lipid standard data, these observations may indicate ISF due to postionization. Finally, while artifact contributions from matrix blanks were also noted, both experimental and computational data sets suggest that the overall extent of ISF is statistically increased in MALDI-2 compared to MALDI-1.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty Estimates for Sonic-Boom Pressure Signatures and Loudness Carpets

A non-intrusive uncertainty quantification method is applied to computational analysis of supersonic, low-boom aircraft. The mean and standard deviation statistics of the pressure waveforms and loudness metrics are evaluated through use of numerical quadrature. The probability density function (p.d.f.) of these outputs is evaluated via kernel density estimation. The simulations use an inviscid, embedded-boundary Cartesian-mesh flow solver in the nearfield combined with an augmented Burgers’ equation solver for propagation in the farfield. The results show that the p.d.f. of the waveform is bimodal at shocks, which makes the mean and standard deviation statistics inappropriate. Despite this limitation, we show that the moment statistics can provide effective assessment of discrepancies when comparing with experimental data. This is demonstrated by presenting uncertainty analysis of a wind-tunnel test and showing that we significantly improve the predictions when we include the test uncertainties in the simulation. Normal distributions are obtained for the ground signature and loudness metrics, which is primarily due to the careful shaping of the low-boom waveform. Separation of variables and error control are used to reduce computational cost. We demonstrate that this is an efficient approach in the sense of balancing numerical errors in the statistics quadrature with discretization errors in the solvers.

ARMD↗

Stochastic multiscale modeling for quantifying statistical and model errors with application to composite materials

This paper provides a coherent and efficient computational framework for stochastic multiscale analysis of material systems in the presence of parametric uncertainties and modeling errors. Uncertainty in those model parameters that are not deduced as upscaled quantities is attributed to an uncertainty “germ”. While such parameters can appear at any scale, they are predominant at the finest analysis scale. Additional uncertainties stemming from statistical estimation, attributed to lack of data and model error, are associated with each submodel contributing to the multiscale system. Here, a robust and efficient framework based on a generalized extended polynomial chaos expansion (gEPCE) is proposed to simultaneously propagate all these uncertainties in order to provide a probabilistic representation of specific quantities of interest (QoI). We characterize the full probability distribution of the QoI and the uncertainty in the failure probability pertaining to its tails. By combining gEPCE with kernel density estimation (KDE) and directional derivatives, we construct sensitivity measures that connect these statistical metrics of QoI to the various sources of uncertainty to assess their individual and combined impacts. An illustrative problem featuring three-point bending of a composite beam is investigated to demonstrate the presented approach.

36 MATERIALS SCIENCE↗

A group finder algorithm optimised for the study of local galaxy environments

Context. The majority of galaxy group catalogues available in the literature use the popular friends-of-friends algorithm which links galaxies using a linking length. One potential drawback to this approach is that clusters of points can be linked with thin bridges which may not be desirable. In order to study galaxy groups, it is important to obtain realistic group structures. Aim. Here, in this study, we present a new simple group finder algorithm, TD-ENCLOSER, that finds the group that encloses a target galaxy of interest. Methods. TD-ENCLOSER is based on the kernel density estimation method which treats each galaxy, represented by a zero-dimensional particle, as a two-dimensional circular Gaussian. The algorithm assigns galaxies to peaks in the density field in order of density in descending order (‘top down’) so that galaxy groups ‘grow’ around the density peaks. Outliers in under-dense regions are prevented from joining groups by a specified hard threshold, while outliers at the group edges are clipped below a soft (blurred) interior density level. Results. The group assignments are largely insensitive to all free parameter variations apart from the hard density threshold and the kernel standard deviation, although this is a known feature of density-based group finder algorithms and it operates with a computing speed that increases linearly with the size of the input sample. In preparation for a companion paper, we also present a simple algorithm to select unique representative groups when duplicates occur. Conclusions. TD-ENCLOSER is tested on a mock galaxy catalogue using a smoothing scale of 0.3 Mpc and is found to be able to recover the input group distribution with sufficient accuracy to be applied to observed galaxy distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning materials properties with accurate predictions, uncertainty estimates, domain guidance, and persistent online accessibility

One compelling vision of the future of materials discovery and design involves the use of machine learning (ML) models to predict materials properties and then rapidly find materials tailored for specific applications. However, realizing this vision requires both providing detailed uncertainty quantification (model prediction errors and domain of applicability) and making models readily usable. At present, it is common practice in the community to assess ML model performance only in terms of prediction accuracy (e.g. mean absolute error), while neglecting detailed uncertainty quantification and robust model accessibility and usability. Here, we demonstrate a practical method for realizing both uncertainty and accessibility features with a large set of models. We develop random forest ML models for 33 materials properties spanning an array of data sources (computational and experimental) and property types (electrical, mechanical, thermodynamic, etc). All models have calibrated ensemble error bars to quantify prediction uncertainty and domain of applicability guidance enabled by kernel-density-estimate-based feature distance measures. All data and models are publicly hosted on the Garden-AI infrastructure, which provides an easy-to-use, persistent interface for model dissemination that permits models to be invoked with only a few lines of Python code. We demonstrate the power of this approach by using our models to conduct a fully ML-based materials discovery exercise to search for new stable, highly active perovskite oxide catalyst materials.

domain of applicability↗

Non-uniform active learning for Gaussian process models with applications to trajectory informed aerodynamic databases

The ability to non-uniformly weight the input space is desirable for many applications, and has been explored for space-filling approaches. Increased interests in linking models, such as in a digital twinning framework, increases the need for sampling emulators where they are most likely to be evaluated. In particular, here we apply non-uniform sampling methods for the construction of aerodynamic databases. This paper combines non-uniform weighting with active learning for Gaussian Processes (GPs) to develop a closed-form solution to a non-uniform active learning criterion. We accomplish this by utilizing a kernel density estimator as the weight function. We demonstrate the need and efficacy of this approach with an atmospheric entry example that accounts for both model uncertainty as well as the practical state space of the vehicle, as determined by forward modeling within the active learning loop.

42 ENGINEERING↗