Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “density estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗

Raccoon density estimation from camera traps for raccoon rabies management

Abstract Density estimation for unmarked animals is particularly challenging, yet density estimates are often necessary for effective wildlife management. Raccoons ( Procyon lotor ) are the primary terrestrial wildlife reservoir for Lyssavirus rabies within the United States. The raccoon rabies variant (RRVV) is actively managed at landscape scales using oral rabies vaccination (ORV) within the eastern United States. To effectively manage RRVV, it is important to know the density of raccoons to appropriately scale the density of ORV baits distributed on the landscape. We compared methods to estimate raccoon densities from camera‐trap data versus more intensive capture‐mark‐recapture (CMR) estimates across 2 land cover types (upland pine and bottomland hardwood) in the southeastern United States during 2019 and 2020. We evaluated the effect of alternative camera configurations and durations of camera trapping on density estimates and used an N‐mixture model to estimate raccoon densities, including covariates on abundance and detection. We further compared different methods of scaling camera‐based counts, with the maximum number of raccoons seen on any given image within a day best explaining density. Camera‐trap density estimates were moderately correlated with CMR estimates ( r = 0.56). However, densities from camera‐trap data were more reliable when classifying category of density as an index used to inform management (83% correct when compared to CMR estimates), although the densities in our study fell into the 2 lowest density classes only. Using more cameras reduced bias and uncertainty around density estimates; however, if ≤6 camera traps were used at a site, a line transect approach proved less biased than a grid design. Camera trapping should be conducted for at least 3 weeks for more accurate estimates of raccoon population density in our study area (<5% bias). We show that camera‐trap data can be used to assign raccoon densities to management‐relevant density index bins, but more studies are needed to ensure reliability across a greater range of environmental conditions and raccoon densities.

Davis, Amy J.↗

Maximum Switching Throughput Density Estimator

SAND2024-11125O The Maximum Switching Throughput Density Estimator software performs a simple analysis that estimates the maximum logic switching throughput density that’s achieved in various CMOS technology nodes on the International Roadmap for Devices and Systems. This software utilizes simple device models and optimization techniques, performing a simple sweep over a range of possible logic supply voltages, and analytically calculating the maximum switching frequency for the given logic voltage that meets the power density constraint. It does this by using simple models of power dissipation in conventional and fully adiabatic switching. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Frank, Michael↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

Nonparametric, data-based kernel interpolation for particle-tracking simulations and kernel density estimation

Traditional interpolation techniques for particle tracking include binning and convolutional formulas that use pre-determined (i.e., closed-form, parameteric) kernels. In many instances, the particles are introduced as point sources in time and space, so the cloud of particles (either in space or time) is a discrete representation of the Green’s function of an underlying PDE. As such, each particle is a sample from the Green’s function; therefore, each particle should be distributed according to the Green’s function. In short, the kernel of a convolutional interpolation of the particle sample “cloud” should be a replica of the cloud itself. This idea gives rise to an iterative method by which the form of the kernel may be discerned in the process of interpolating the Green’s function. When the Green’s function is a density, this method is broadly applicable to interpolating a kernel density estimate based on random data drawn from a single distribution. We formulate and construct the algorithm and demonstrate its ability to perform kernel density estimation of skewed and/or heavy-tailed data including breakthrough curves.

42 ENGINEERING↗

Diffusion-Model-Assisted Supervised Learning of Generative Models for Density Estimation

Here, we present a supervised learning framework of training generative models for density estimation. Generative models, including generative adversarial networks (GANs), normalizing flows, and variational auto-encoders (VAEs), are usually considered as unsupervised learning models, because labeled data are usually unavailable for training. Despite the success of the generative models, there are several issues with the unsupervised training, e.g., requirement of reversible architectures, vanishing gradients, and training instability. To enable supervised learning in generative models, we utilize the score-based diffusion model to generate labeled data. Unlike existing diffusion models that train neural networks to learn the score function, we develop a training-free score estimation method. This approach uses mini-batch-based Monte Carlo estimators to directly approximate the score function at any spatial-temporal location in solving an ordinary differential equation (ODE), corresponding to the reverse-time stochastic differential equation (SDE). This approach can offer both high accuracy and substantial time savings in neural network training. Once the labeled data are generated, we can train a simple, fully connected neural network to learn the generative model in the supervised manner. Compared with existing normalizing flow models, our method does not require the use of reversible neural networks and avoids the computation of the Jacobian matrix. Compared with existing diffusion models, our method does not need to solve the reverse-time SDE to generate new samples. As a result, the sampling efficiency is significantly improved. We demonstrate the performance of our method by applying it to a set of 2D datasets as well as real data from the University of California Irvine (UCI) repository.

97 MATHEMATICS AND COMPUTING↗

The effective number of parameters in kernel density estimation

We devise a new formula for measuring the effective degrees of freedom (EDoF) in kernel density estimation (KDE). Starting from the orthogonal polynomial sequence (OPS) expansion for the ratio of the empirical to the oracle density, we show how convolution with the kernel leads to a new OPS with respect to which one may express the resulting KDE. The expansion coefficients of the two OPS systems can then be related via a kernel sensitivity matrix, which leads to a natural oracle definition of EDoF through the trace operator. Asymptotic properties of the (empirical) plug-in EDoF are worked out through influence functions, and connections with other empirical EDoFs are established. Minimization of Kullback-Leibler divergence is investigated as an alternative to integrated squared error based bandwidth selection rules, yielding a new normal scale rule. The methodology, which arises from a proper oracle formulation and is not restricted to convolution kernels, suggests the possibility of a new bandwidth selection rule based on an information criterion such as AIC.

bandwidth selection↗

Photometric redshifts probability density estimation from recurrent neural networks in the DECam local volume exploration survey data release 2

Photometric wide-field surveys are imaging the sky in unprecedented detail. These surveys face a significant challenge in efficiently estimating galactic photometric redshifts while accurately quantifying associated uncertainties. In this work, we address this challenge by exploring the estimation of Probability Density Functions (PDFs) for the photometric redshifts of galaxies across a vast area of 17,000 square degrees, encompassing objects with a median 5 σ point-source depth of g = 24.3, r = 23 . 9 , i = 23.5, and z = 22.8 mag. Our approach uses deep learning, specifically integrating a Recurrent Neural Network architecture with a Mixture Density Network, to leverage magnitudes and colors as input features for constructing photometric redshift PDFs across the whole DECam Local Volume Exploration (DELVE) survey sky footprint. Subsequently, we rigorously evaluate the reliability and robustness of our estimation methodology, gauging its performance against other well-established machine learning methods to ensure the quality of our redshift estimations. Our best results constrain photometric redshifts with the bias of − 0 . 0013 , a scatter of 0.0293, and an outlier fraction of 5.1%. These point estimates are accompanied by well-calibrated PDFs evaluated using diagnostic tools such as Probability Integral Transform and Odds distribution. We also address the problem of the accessibility of PDFs in terms of disk space storage and the time demand required to generate their corresponding parameters.We present a novel Autoencoder model that reduces the size of PDF parameter arrays to one-sixth of their original length, significantly decreasing the time required for PDF generation to one-eighth of the time needed when generating PDFs directly from the magnitudes.

79 ASTRONOMY AND ASTROPHYSICS↗

Plasma Properties in the Earth's Magnetosheath Near the Subsolar Magnetopause: Implications for Geocoronal Density Estimates

Combined in situ ion measurements and remote sensing of energetic neutral atoms are used to determine the geocoronal Hydrogen density at large (∼10 R E ) distances from the Earth. This method for determining the geocoronal density requires global magnetospheric modeling. Observations in the Earth's subsolar magnetosheath from the Magnetospheric Multiscale mission are used to determine the accuracy of using global models to predict the geocoronal density. On average, gas dynamic and magnetohydrodynamic (MHD) models and observations are in reasonable agreement, with differences <25%. In addition, the MHD model subsolar magnetopause is about 0.5 R E sunward of the observed location. However, variations around averages are large (up to a factor of 2), indicating that global models introduce relatively large uncertainties in geocoronal density estimates. Finally, the critical ion flux in the Interstellar Boundary Explorer IBEX‐Hi energy range is often minimally affected by fluctuations of a factor of 2 in the density.

79 ASTRONOMY AND ASTROPHYSICS↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Sampling Size Optimization for Bioburden Density Estimation in Planetary Protection

Planetary protection (PP) is a discipline that focuses on minimizing the biological contamination of spacecraft to ensure compliance with international policy. Precise estimation of bioburden - the total number of microbes in or on spacecraft hardware – and the bioburden density are of utmost importance for PP. Such estimation is the way concordance with requirements is demonstrated, and it is critical for quantifying the potential risk of inadvertently contaminating other planetary bodies. Although a suite of molecular techniques have been used to thoroughly characterize and profile the microbiome of various cleanroom environments and spacecraft, the gold standard remains the physical enumeration of microbes via culturing of samples directly taken from spacecraft and associated surfaces. However, due to technical, budgetary, and programmatic constraints, only a manageable portion (around 10%) of the entire spacecraft surface is directly sampled with cotton swabs or wipes. To generate the bioburden current best estimate (CBE) for components not directly verifiable, the accepted approach is to apply a NASA-defined bioburden estimate based on the components’ manufacturing or assembly environment. This approach utilizes a prespecified bioburden density estimation that applies a maximum value across the total surface area of the specified component. For hardware components that underwent similar assembly processes, an implied bioburden is adopted for all components, based on a direct verification of a representative component within the same lot. Once all components have a CBE, the bioburden estimates are generated. In previous publication [ 1], we have shown that statistical risks quantifying the accuracy of the estimates for sampled, prespecified, and implied components can be derived and ranked. For mean squared error (MSE) function, the risks are available analytically and hence a cost function can be obtained to optimize the risks with respect to the sampling area and sampling cost. Since the sampling area and sampling cost are two complimentary variables, their sum will have a well-defined minimum. This paper presents the multivariate optimization of the integrated risk of an empirical Bayes estimator to determine the optimal sampling schedule for a given number of components. It is assumed that given a number of components, N, the bioburden density for each component can either be sampled, implied, or prespecified. The multivariate optimization searches through different options to sample, imply or prespecify the bioburden density for a component, and account for the component’s surface area and cost of sampling. The idea of the optimization is based on the observation that the statistical risk of using an estimator is a monotonically decreasing function of the sampled area. The larger the sampled area, the lower the risk of using the estimator as the estimator becomes more and more accurate as the sampling area increases. On the other hand, the cost of sampling is monotonically increasing as the sampled surface grows. This makes the risk and total cost of sampling complimentary variables which can be counterbalanced to achieve an optimal overall value with respect to the sampled surface. In this paper, the integrated risk has been used to quantify the accuracy of the estimator. This risk has been selected because it depends on neither the true value of the parameter nor on the collected data. The cost of each sample was also available to obtain the total cost of sampling of N components. The paper will present the results based on computer-simulated data as well as the data collected during the InSight mission. The computer-simulated data have N components with randomly generated total areas and each component assigned to one of the three categories according to the method of estimating of bioburden density: sampled, implied, or prespecified. The cost of sampling is also available. The cost of sampling is estimated based on a cost model provided by the planetary protection group at JPL. For this paper, the overall cost was assumed to be a linear function of exposure. The optimization process finds the allocation of the components to the three categories that minimizes the tradeoff between integrated risk and total cost. For the InSight data, a set of components is selected representing all three categories, and optimization is performed to determine if the performed allocation was optimal or if a better allocation could have been obtained. To the best of our knowledge, this work is the first attempt not only perform an accurate estimation of bioburden density but also do it in an optimal way.

97 - MATHEMATICS AND COMPUTING↗

Out-of-distribution detection with non-parametric density estimation for models predicting processing history of uranium ore concentrates

The rapid advancement in machine learning (ML) and computer vision (CV) coincides with the growth of interest in deploying these ML/CV models in numerous fields from medicine to social science. Similar to those areas, we have witnessed a great number of works in materials science employing ML/CV models – neural networks in particular – in their studies in recent years. These models have proven to obtain accurate performance in various tasks. However, these models struggle to attain a similar performance when encountering test samples coming from a distribution that is different from the training set. More importantly, they fail without providing any warning to the users. Therefore, we propose a framework for detecting out-of-distribution (OOD) samples to alert users when a human intervention might be necessary in this work. Specifically, we explore the use of a non-parametric density estimation method to detect OOD samples. Here, we assess OOD detection capability of the proposed framework on ML models developed for categorizing precipitation routes of U 3 O 8 when encountering OOD datasets that contain samples (1) undergone different imaging acquisition process, (2) undergone different material synthesis process, and (3) different materials than ID set. Through those experiments, we achieve an average area under the receiver operating characteristic (AUROC) of at least 91% on average in detecting OOD samples. With minimal overhead cost and superior performance, the proposed framework enables a reliable and safe system when deploying in real-world scenarios.

Convolutional neural networks↗

Spectral-density estimation with the Gaussian integral transform

The spectral-density operator $\hat{ρ}(ω) = δ(ω–\hat{H})$ plays a central role in linear response theory as its expectation value, the dynamical response function, can be used to compute scattering cross sections. In this work, we describe a near optimal quantum algorithm providing an approximation to the spectral density with energy resolution $\Delta$ and error $\epsilon$ using $O(\sqrt{\text{log}_2 (1/ε)[\text{log}_2 (1 / Δ) + \text{log}_2 (1/ε)]/ Δ)}$ operations. This is achieved without using expensive approximations to the time-evolution operator, but instead exploiting qubitization to implement an approximate Gaussian integral transform of the spectral density. Finally, we also describe appropriate error metrics to assess the quality of the spectral function approximations more generally.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bayesian Framework for Bioburden Density Estimation in Planetary Protection

To comply with the international planetary protection policy set forth by the Committee on Space Research and NASA Agency level requirements, spacecraft destined to biologically sensitive planetary bodies have to minimize terrestrial biological contamination. Analysis, testing and inspection are the standard forward verification activities that are used to demonstrate compliance with the biological contamination requirements. For testing of spacecraft surface areas, a swab or wipe sample is collected from surfaces prior to last access and subsequently processed in the lab using NASA Approved Planetary Protection Methods for Culture Based Assays. Raw data resulting from this assay is then statistically treated employing a mathematical paradigm stemming from the 1970’s Viking Lander Project to generate the bioburden density and total microbial bioburden present. This standard approach arbitrarily accounts for error and provides an upper conservative bound as it reports the maximum number of spores estimated to be present on flight hardware surfaces. A bioburden density estimate factors in the following variables: the observed bioburden count, representative volume processed, sampling efficiencies. Notably, to account for error in the approach, a 0 observed count is arbitrarily changed to a count of 1 for each hardware grouping. The data generated by spacecraft bioburden verification campaigns in the past have resulted in <80% of wipes and <90% of swabs containing a bioburden count of 0. As such, having a robust and well documented statistical approach for dealing with the probability of low incident rates is necessary to be able to estimate spacecraft bioburden. Being able to statistically describe the bioburden distribution and associated confidence level is a gamechanger for the development of bioburden allocations during mission design and will allow for tighter management of risk throughout spacecraft build. Thus, Empirical Bayes statistical approach was evaluated to estimate the microbial bioburden on spacecraft to mitigate the aforementioned mathematical concerns and provide a probabilistic bioburden distribution of the flight hardware surface. For application of this approach to performing bioburden calculations, a range of non-informative prior assumptions on hardware surfaces are explored for Bayesian analyses while informative priors using posterior distributions from prior assays are utilized for Empirical Bayes analyses. Several non-informative priors are currently under investigation to assess fitness including use of these priors to serve as a foundation to build off of NASA specification values or a basis of risk to account for unknowns during the integration and testing process. Informative priors under consideration are generated using sampled bioburden values from hardware originating within like processing environments (e.g. vendor cleaning process or similar assembly process), temporal spacecraft status events as a prediction for hardware cleanliness of future samples, and heritage system bioburden actuals to predict allocation for subsequent missions. Informative priors and probabilistic bioburden distributions are then validated using data sets from the Mars Exploration Rover, Mars Science Laboratory, and InSight missions. Using Empirical Bayes approach to generate a probabilistic bioburden distribution as demonstrated through mission use cases provides a valid approach for use in the end-to-end requirements verification process.

97 - MATHEMATICS AND COMPUTING↗

Comparative analysis of plasticity-based GND density estimation methods in crystal plasticity finite element models

In crystal plasticity finite element (CPFE) simulations, accurately quantifying geometrically necessary dislocations (GNDs) is critical for capturing strain gradients in polycrystals. We compare different methods for quantifying GNDs, all of which originate from the Nye tensor, which is computed as the curl of the plastic deformation gradient. The projection technique directly decomposes the Nye tensor onto individual screw and edge dislocation components to compute GNDs. This approach requires converting a nine-component Nye tensor into densities for a larger number of dislocation systems, a fundamentally underdetermined (non-unique) process, which is resolved using L2 minimization. In contrast, when employing CPFE analysis, one could directly compute dislocation densities on each slip system using shear gradients. Projection and slip gradient methods are compared with respect to their prediction of GNDs with changing grain size, strain, and grain neighborhoods, including multigrain junctions. Although these techniques match analytical GND densities for single slip, single crystal deformation, and are consistent with anticipated overall GND trends, we find that the GND densities from projection techniques are significantly lower than those predicted from CPFE-based slip gradients in polycrystals. A suggested improvement of only using the active dislocation systems in the projection technique almost entirely resolved this mismatch.

Crystal plasticity↗

An efficient computational framework for charge density estimation in twisted bilayer graphene

Electronic properties such as band structure and Fermi velocity in low-angle twisted bilayer graphene (TBG) are intrinsically dependent on the atomic structure. Rigid rotation between individual graphene layers provides an approximate description of the bilayer symmetry. Upon relaxation, in-plane displacement of the atoms in low angle TBG causes a change in the symmetry through the enlargement of the AB stacking regions and the reduction in size of AA and SP stacking regions. However, the effect of this in-plane relaxation on the charge density remains unexplored, because the necessary electronic structure calculations of such large supercells of low twist angle TBG are computationally infeasible. Therefore, we develop a computationally efficient framework that enables the exploration of the charge density symmetry of the low twist angle TBG. This framework is based on the Fourier representation of the charge density which presents high intensity Bragg peaks. Here we find that with the decrease of twist angle, low intensity satellite peaks also become apparent. Our framework incorporates these satellite peaks which reveals transformation of symmetry in the charge density distribution from high to low twist angle TBG. One striking outcome is the demonstration of the electron localization in the AA region of low twist angle TBG. Our framework helps to explain the effect of the atomistic relaxation on the charge density distribution and thus, it provides information about exotic electronic properties of low twist angle TBG at a low computational expense.

36 MATERIALS SCIENCE↗