Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel density estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A meshless stochastic method for Poisson–Nernst–Planck equations

A plethora of biological, physical, and chemical phenomena involve transport of charged particles (ions). Its continuum-scale description relies on the Poisson–Nernst–Planck (PNP) system, which encapsulates the conservation of mass and charge. The numerical solution of these coupled partial differential equations is challenging and suffers from both the curse of dimensionality and difficulty in efficiently parallelizing. We present a novel particle-based framework to solve the full PNP system by simulating a drift–diffusion process with time- and space-varying drift. We leverage Green’s functions, kernel-independent fast multipole methods, and kernel density estimation to solve the PNP system in a meshless manner, capable of handling discontinuous initial states. The method is embarrassingly parallel, and the computational cost scales linearly with the number of particles and dimension. We use a series of numerical experiments to demonstrate both the method’s convergence with respect to the number of particles and computational cost vis-à-vis a traditional partial differential equation solver.

Chemistry↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

Polynomial Chaos Surrogate Construction for Random Fields with Parametric Uncertainty

Engineering and applied science rely on computational experiments to rigorously study physical systems. The mathematical models used to probe these systems are highly complex, and sampling-intensive studies often require prohibitively many simulations for acceptable accuracy. Surrogate models provide a means of circumventing the high computational expense of sampling such complex models. In particular, polynomial chaos expansions (PCEs) have been successfully used for uncertainty quantification studies of deterministic models where the dominant source of uncertainty is parametric. We discuss an extension to conventional PCE surrogate modeling to enable surrogate construction for stochastic computational models that have intrinsic noise in addition to parametric uncertainty. We develop a PCE surrogate on a joint space of intrinsic and parametric uncertainty, enabled by Rosenblatt transformations, which are evaluated via kernel density estimation of the associated conditional cumulative distributions. Furthermore, we extend the construction to random field data via the Karhunen–Loève expansion. We then take advantage of closed-form solutions for computing PCE Sobol indices to perform a global sensitivity analysis of the model which quantifies the intrinsic noise contribution to the overall model output variance. Additionally, the resulting joint PCE is generative in the sense that it allows generating random realizations at any input parameter setting that are statistically approximately equivalent to realizations from the underlying stochastic model. The method is demonstrated on a chemical catalysis example model and a synthetic example controlled by a parameter that enables a switch from unimodal to bimodal response distributions.

97 MATHEMATICS AND COMPUTING↗

PV-Finder: ML Based Algorithm for Primary Vertex Identification

he CMS detector at the High-Luminosity Large Hadron Collider (HL-LHC) will operate in challenging conditions with expected pile-up of up to 200 collisions per bunch crossing, necessitating the development of a more resilient primary vertex (PV) reconstruction method to ensure the integrity of data analysis and the efficiency of the CMS triggering system. This contribution describes preliminary studies on a new ML based PV-Finder method for PV identification. The method is based on a model trained using Kernel Density Estimations (KDEs) derived from the positions of reconstructed tracks at the beamline, incorporating uncertainties from track parameters. It also utilizes target histograms, modeled as Gaussian distributions centered on the actual ground truth values of specific primary vertices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Non-Parametric Statistical Analysis of Current Waveforms through Power System Sensors

The protection, control, and monitoring of the power grid is not possible without accurate measurement devices. As the percentage of renewable energy sources penetrating the existing grid infrastructure increases, so do uncertainties surrounding their effects on the everyday operation of the power system. Many of these devices are sources of high-frequency transients. These transients may be useful for identifying certain events or behaviors otherwise not seen in traditional analysis techniques. Therefore, the ability of sensors to accurately capture these phenomena is paramount. In this work, two commercial-grade power system distribution sensors are investigated in terms of their ability to replicate high-frequency phenomena by studying their responses to three events: a current inrush, a microgrid “close-in”, and a fault on the terminals of a wind turbine. Kernel density estimation is used to derive the non-parametric probability density functions of these error distributions and their adequateness is quantified utilizing the commonly used root mean square error (RMSE) metric. It is demonstrated that both sensors exhibit characteristics in the high harmonic range that go against the assumption that measurement error is normally distributed.

47 OTHER INSTRUMENTATION↗

Improvement of the NOvA Near Detector Event Reconstruction and Primary Vertexing through the Application of Machine Learning Methods

The purpose of this work is to examine the application of a deep learning model in event reconstruction of neutrino interactions. The challenges faced in event reconstruction include the placement of an accurate primary neutrino interaction vertex which is used to support the particle track and prong algorithms. The result of accurate primary vertex ensures all particles involved in a neutrino interaction are included. We propose a regression-based Convolutional Neural Network (CNN) method to predict the primary vertex of a particle interaction. We show that with raw two-dimensional pixel map views as input, the regression-based CNN can predict the primary vertex in all three coordinates. This work is applied as part of the NOvA (NuMI Off-axis $\nu_e$ Appearance) near detector reconstruction efforts. The primary vertex predicted by the regression-based CNN model shows promising results for future applications. This deep learning method can be extended to secondary vertexing through a Kernel Density Estimate algorithm discussed in this work.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A VOI Web Application for Distinct Geothermal Domains: Statistical Evaluation of Different Data Types within the Great Basin

The Great Basin region contains different domains that have different structural and hydrothermal flow patterns. Depending on the characteristics of these patterns, certain data types may be more successful at detecting hidden geothermal resources. In this paper, we quantitatively evaluate if certain data types are more successful in certain domains. Given different aquifer, strain and structural conditions, we explore which data types statistically reveal positively labeled geothermal sites. We utilize value of information (VOI) metrics to help quantify the reliability of data types to discriminate against "positive" and "negative" labeled geothermal sites. We also evaluate how kernel density estimation can help generalize the statistics that inform VOI, which is necessary given the limited data in geothermal exploration. Except for the Carbonate Aquifer, the highest ranking of the Vimperfect is the Local Structural Setting. Next, the slip and dilation tendency is first for Carbonate Aquifer and second for Central Nevada Seismic Belt and Western Great Basin. For the Carbonate Aquifer, heat flow is has the lowest Vimperfect value compared to the other three domains, which is consistent with the understanding of how heat flow measurements are masked by regional groundwater flow.

Bayesian analysis↗

Minimum entropy filtering for a single output non-Gaussian stochastic system using state transformation

This paper presents a novel filter design for the single-output stochastic non-linear systems subjected to non-Gaussian noises and the proposed assumptions. Based on a state transformation, the unmeasurable states of the systems can be estimated where non-linear terms in the systems have been eliminated. It has been shown that the estimation error is linearly dynamical regarding to the presented vector-valued filter gain which can be optimised by minimising the entropy-based performance criterion. In addition, the convergence of the presented algorithm is analysed in mean-square sense and a numerical example is given to verify the effectiveness of the presented filtering algorithm. Meanwhile, the extended Kalman filter, unscented particle filter and minimum entropy filter are given for the comparisons of the filtering performance. Following the presented framework, some extensions of the presented filtering algorithm are discussed to indicate the flexibility of the filter design. The contribution of this paper can be summarised as establishing a novel minimum entropy filtering framework which consists of model transformation, entropy optimisation and convergence analysis.

42 ENGINEERING↗

On the measurement of shape: With applications to lunar regolith

With the renewed commitment from NASA and other commercial entities for a presence on the Moon, the importance of understanding the characteristics of lunar regolith and how to utilize it have become the target of increasing scrutiny. Much of what is known about lunar regolith was collected during and immediately after the Apollo program, however, analytical techniques and instrumentation have advanced in leaps and bounds in the subsequent decades. Specifically, dynamic image analysis systems have advanced to the point that millions of particles can have morphological characteristics automatically determined in relatively short time frames. Particle morphological data was collected on several lunar samples and a pair of widely used regolith simulants to ascertain the accuracy of these simulants and to explore statistical analysis methods of these large datasets. It is found that these morphology datasets can vary widely depending on the particle size of the particles, and simple averaging of the data skews the results heavily towards the numerically abundant size ranges, the fines. Different reporting methods are suggested to ameliorate these problems. When applied to the lunar regolith, the particles are noted to be less morphologically complex than initially suspected. Compared to the lunar material, the simulants are found to contain some more morphological variability and angular grains. Such difference is likely due to the wildly different comminution processes that these different powder systems are subjected to.

2D shape↗

Data-driven Minimum Entropy Control for Stochastic Nonlinear Systems using the Cumulant-Generating Function

Here, we present a novel minimum entropy control algorithm for a class of stochastic nonlinear systems subjected to non-Gaussian noises. The entropy control can be considered as an optimization problem for the system randomness attenuation, but the mean value has to be considered separately. To overcome this disadvantage, a new representation of the system stochastic properties was given using the cumulant-generating function based on the moment-generating function, in which the mean value and the entropy was reflected by the shape of the cumulant-generating function. Based on the samples of the system output and control input, a time-variant linear model was identified, and the minimum entropy optimization was transformed to system stabilization. Then, an optimal control strategy was developed to achieve the randomness attenuation, and the boundedness of the controlled system output was analyzed. The effectiveness of the presented control algorithm was demonstrated by a numerical example. In this paper, a data-driven minimum entropy design is presented without pre-knowledge of the system model; entropy optimization is achieved by the system stabilization approach in which the stochastic distribution control and minimum entropy are unified using the same identified structure; and a potential framework is obtained since all the existing system stabilization methods can be adopted to achieve the minimum entropy objective.

42 ENGINEERING↗

Spatial and temporal overlap between hatchery- and natural-origin steelhead and Chinook Salmon during spawning in the Klickitat River, Washington, USA

Abstract Objective A goal of many segregated salmonid hatchery programs is to minimize potential interbreeding between hatchery- and natural-origin fish. Our objective was to assess this on the Klickitat River, Washington, USA. Methods We used radiotelemetry to evaluate spatiotemporal spawning overlap between hatchery- and natural-origin steelhead Oncorhynchus mykiss and spring Chinook Salmon O. tshawytscha. We estimated percentages of tagged fish that spawned naturally in the Klickitat River subbasin, emigrated from the Klickitat River, or died before spawning. A kernel density analysis was used to estimate probability of spatiotemporal overlap between hatchery- and natural-origin spawners. Result For steelhead, 12% of hatchery-origin and 50% of natural-origin fish spawned naturally. For spring Chinook Salmon, 18% of hatchery-origin and 44% of natural-origin fish spawned naturally. Tag loss may result in underestimates in these percentages. Most hatchery-origin steelhead (90%) spawned downstream of river kilometer (rkm) 32, and 75% spawned from November to mid-March. The majority of natural-origin steelhead (64%) spawned upstream of rkm 32, and 75% spawned from mid-March to late May. Spawn timing of hatchery-origin Chinook Salmon (early August to mid-September) overlapped with that of natural-origin Chinook Salmon (late July to late September), and fish of both origins spawned in the same 30-km reach of the river. We estimated the percentage of hatchery-origin spawners (pHOS) on the natural spawning grounds to be 12% for steelhead and 40% for spring Chinook Salmon across all study years. For steelhead, we estimated the overlap probability to be 25% (95% CI = 22.5–28%). For spring Chinook Salmon, tight spatial clustering of hatchery-origin fish resulted in a lower overlap estimate of 21% (13–31%). Conclusion We suggest adjusting pHOS estimates using these overlap estimates or similar spatiotemporal data on actual spawner proximity and possible interactions, and that these types of analyses be used in conjunction with gene flow analysis to accurately evaluate effects of individual hatchery programs.

Zendt, Joseph S.↗

Computing the QRPA level density with the finite amplitude method

Here, we describe a new algorithm to calculate the vibrational nuclear level density of an atomic nucleus. Fictitious perturbation operators that probe the response of the system are generated by drawing their matrix elements from some probability distribution function. We use the Finite Amplitude Method to explicitly compute the response for each such sample. With the help of the Kernel Polynomial Method, we build an estimator of the vibrational level density and provide the upper bound of the relative error in the limit of infinitely many random samples. The new algorithm can give accurate estimates of the vibrational level density. Since it is based on drawing multiple samples of perturbation operators, its computational implementation is naturally parallel and scales like the number of available processing units.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

New Machine Learning Techniques for Simulation-Based Inference: InferoStatic Nets, Kernel Score Estimation, and Kernel Likelihood Ratio Estimation

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential $\varphi$. In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

New machine learning techniques for simulation-based inference: InferoStatic nets, kernel score estimation, and kernel likelihood ratio estimation

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

Kong, Kyoungchul↗

Training quantum neural networks using the quantum information bottleneck method

Abstract We provide in this paper a concrete method for training a quantum neural network to maximize the relevant information about a property that is transmitted through the network. This is significant because it gives an operationally well founded quantity to optimize when training autoencoders for problems where the inputs and outputs are fully quantum. We provide a rigorous algorithm for computing the value of the quantum information bottleneck quantity within error ε that requires O ( log 2 ⁡ ( 1 / ϵ ) + 1 / δ 2 ) queries to a purification of the input density operator if its spectrum is supported on { 0 } ⋃ [ δ , 1 − δ ] for δ > 0 and the kernels of the relevant density matrices are disjoint. We further provide algorithms for estimating the derivatives of the QIB function, showing that quantum neural networks can be trained efficiently using the QIB quantity given that the number of gradient steps required is polynomial.

Çatlı, Ahmet Burak (ORCID:0000000152294141)↗

Coarse-Grained Density Functional Theory Predictions via Deep Kernel Learning

Scalable electronic predictions are critical for soft materials design. Recently, the Electronic Coarse-Graining (ECG) method was introduced to renormalize all-atom quantum chemical (QC) predictions to coarse-grained (CG) resolutions using deep neural networks (DNNs). While DNNs can learn complex representations that prove challenging for kernel-based methods, they are susceptible to overfitting and the overconfidence of uncertainty estimations. Here, we develop ECG within a GPU-accelerated Deep Kernel Learning (DKL) framework to enable CG QC predictions using range-separated hybrid density functional theory (DFT), obtaining a 107 speedup relative to naive all-atom QC. By treating the predicted electronic properties as random Gaussian Processes, DKL incorporates CG mapping degeneracy by learning the distribution of electronic energies as a function of CG configuration. DKL-ECG accurately reproduces molecular orbital energies from range-separated DFT while facilitating efficient training via active learning using the uncertainties provided by DKL. Further, we show that while active learning algorithms enable efficient sampling of a more diverse configurational space relative to random sampling, all explored query methods exhibit comparable performance for the examined system. We attribute this result to the significant overlap of the feature space and output property distributions across multiple temperatures.

97 MATHEMATICS AND COMPUTING↗