Engineering PapersSearch

SEARCH · Engineering Papers

Results for “count data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Poisson Log-Normal Process for Count Data Prediction

Modeling count data is important in physics and other scientific disciplines, where measurements often involve discrete, non-negative quantities such as photon or neutrino detection events. Traditional parametric approaches can be trained to generate integer-count predictions but may struggle with capturing complex, non-linear dependencies often observed in the data. Gaussian process (GP) regression provides a robust non-parametric alternative to modeling continuous data; however, it cannot generate integer outputs. We propose the Poisson Log-Normal (PoLoN) process, a framework that employs GP to model Poisson log-rates. As in GP regression, our approach relies on the correlations between data points captured via GP kernel structure rather than explicit functional parameterizations. We demonstrate that the PoLoN predictive distribution is Poisson-LogNormal and provide an algorithm for optimizing kernel hyperparameters. Furthermore, we adapt the PoLoN approach to the problem of detecting weak localized signals superimposed on a smoothly varying background - a task of considerable interest in many areas of science and engineering. Our framework allows us to predict the strength, location and width of the detected signals. We evaluate PoLoN's performance using both synthetic and real-world datasets, including the open dataset from CERN which was used to detect the Higgs boson at the Large Hadron Collider. Our results indicate that the PoLoN process can be used as a non-parametric alternative for analyzing, predicting, and extracting signals from integer-valued data.

Saha, Anushka [Rutgers U., Piscataway]

Tensor decompositions for count data that leverage stochastic and deterministic optimization

There is growing interest to extend low-rank matrix decompositions to multi-way arrays, or tensors. One fundamental low-rank tensor decomposition is the canonical polyadic decomposition (CPD). The challenge of fitting a low-rank, nonnegative CPD model to Poisson-distributed count data is of particular interest. Several popular algorithms use local search methods to approximate the maximum likelihood estimator (MLE) of the Poisson CPD model. Here, this work presents two new algorithms that extend state-of-the-art local methods for Poisson CPD. Hybrid GCP-CPAPR combines Generalized Canonical Decomposition (GCP) with stochastic optimization and CP Alternating Poisson Regression (CPAPR), a deterministic algorithm, to increase the probability of converging to the MLE over either method used alone. Restarted CPAPR with SVDrop uses a heuristic based on the singular values of the CPD model unfoldings to identify convergence toward optimizers that are not the MLE and restarts within the feasible domain of the optimization problem, thus reducing overall computational cost when using a multi-start strategy. We provide empirical evidence that indicates our approaches outperform existing methods with respect to converging to the Poisson CPD MLE.

CPAPR

Detecting outbreaks using a spatial latent field

In this paper, we present a method for estimating the infection-rate of a disease as a spatial-temporal field. Our data comprises time-series case-counts of symptomatic patients in various areal units of a region. We extend an epidemiological model, originally designed for a single areal unit, to accommodate multiple units. The field estimation is framed within a Bayesian context, utilizing a parameterized Gaussian random field as a spatial prior. We apply an adaptive Markov chain Monte Carlo method to sample the posterior distribution of the model parameters condition on COVID-19 case-count data from three adjacent counties in New Mexico, USA. Our results suggest that the correlation between epidemiological dynamics in neighboring regions helps regularize estimations in areas with high variance (i.e., poor quality) data. Using the calibrated epidemic model, we forecast the infection-rate over each areal unit and develop a simple anomaly detector to signal new epidemic waves. Our findings show that anomaly detector based on estimated infection-rates outperforms a conventional algorithm that relies solely on case-counts.

Safta, Cosmin [Sandia National Laboratories (SNL-C

Low energy neutron light output characterization of EJ301D and deuterated stilbene with a comparison of light output characterization methods

The neutron-induced light yield of a 2.54 cm diameter by 2.54 cm long right circular cylinder of EJ301D and a (5.08 cm)3 custom made cube of deuterated trans-stilbene-d12 (d-stilbene) were measured over incident neutron energies from 300 keV to 2.2 MeV and 200 keV to 2.4 MeV, respectively. The measurements were performed using a time-of-flight experiment with a Cf-252 source and an approximately 1.5 m flight path. We compare three light output spectrum full energy deposition edge estimation methods: (1) simulating the neutron energy spectrum edge and fitting it to the light output spectrum, (2) using the inflection point of the light output spectrum edge (derivative method, a.k.a. Kornilov’s method), and (3) using an empirical model fit to the edge of the light output spectrum. Both the derivative and equation fit methods do not account for physical processes such as multiple neutron scattering in the detectors. They instead rely on assumptions about the linear shape continuum shape of the light output spectrum and the direct correlation between the location of the spectrum’s inflection point and maximum energy deposition. These assumptions were found to introduce bias into those methods when tested against simulated spectra with known edge locations. When tested against measured spectra the derivative method was found to differ from the simulation fit by greater than 30% at low energies with large discontinuities for adjacent data points above 800 keV incident neutron energy. The empirical equation fitting method was found to also exhibit bias of a similar magnitude, but with significantly more continuous behavior, especially with the lower count data of the smaller volumed EJ301D scintillator. Experimental light output yield for this neutron energy range is reported using the simulated spectrum fitting method because it includes physics neglected by the other methods, and did not exhibit the bias observed in the other methods

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Shot-noise-induced lower temperature limit of the nonneutral plasma parallel temperature diagnostic

Abstract We develop a new algorithm to estimate the temperature of a nonneutral plasma in a Penning-Malmberg trap. The algorithm analyzes data obtained by slowly lowering a voltage that confines one end of the plasma and collecting escaping charges, and is a maximum likelihood estimator based on a physically-motivated model of the escape protocol presented in (Beck in Measurement of the magnetic and temperature dependence of the electron-electron anisotropic temperature relaxation rate. PhD thesis, 1990). Significantly, our algorithm may be used on single-count data, allowing for improved fits with low numbers of escaping electrons. This is important for low-temperature plasmas such as those used in antihydrogen trapping. We perform a Monte Carlo simulation of our algorithm, and assess its robustness to intrinsic shot noise and external noise. The assumptions in this paper allow for a lower bound for measurable plasma temperatures of approximately $3\,\mathrm{K}$ 3 K for plasmas of length $1\,\mathrm{cm}$ 1 cm , with approximately 100 particle counts needed for an accuracy of $\pm 10 \%$ ± 10 % .

Zhong, Adrianne (ORCID:0000000162618736)

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie

Quantifying health benefits of sustainable aviation fuels: Modeling decreased ultrafine particle emissions and associated impacts on communities near the Seattle-Tacoma International Airport

Exposure to ultrafine particles (UFP, ≤100 nm) is an emerging health concern linked to premature mortality, with jet fuel combustion identified as a significant source of UFPs near airports. Sustainable aviation fuel (SAF) adoption has the potential to reduce aviation-related UFPs and may particularly benefit populations who reside nearby. However, assessing aviation-specific impacts on health remains challenging due to the lack of tools capable of addressing: fine-scale exposure evaluation, novel ambient pollutants, and groups with increased exposure or susceptibility. We develop and apply a method to estimate reductions in mortality associated with aviation-related UFP reductions at the Seattle-Tacoma (SEA-TAC) International Airport under SAF adoption scenarios, with a focus on near-airport communities. Using UFP exposure surfaces generated from AERMOD modeling, flight count data, and UFP measurements, we evaluated UFP reductions under various control scenarios. We estimated mortality reductions by combining this with population data, baseline mortality, and a hazard ratio of 1.012 (95 % confidence interval: 1.010, 1.015) per interquartile range increment of 2723 particles/cm 3 . Our analysis included 412 census tracts representing almost 1.5 million adults. Baseline aviation-related UFP exposures averaged 1145 (SD: 277) particles/cm 3 . The highest baseline concentrations and subsequent reductions under SAF scenarios were near SEA-TAC. Mortality case reductions averaged between 3.1 (95 % range: 2.5–3.7) for a 5 % UFP reduction to 31.0 (24.6–37.4) for a 50 % reduction, with corresponding mortality rate reductions of 0.2 (0.2–0.3) to 2.1 (1.7–2.5) cases per 100,000 people per year. Mortality rate reductions were larger among populations residing closer to SEA-TAC, including those that were Hispanic or Latino, below-poverty, and did not identify as White. Reducing aviation-related UFPs through SAF adoption could lead to lower mortality, particularly in near-airport communities. This reproducible approach can be adapted to other settings to evaluate health benefits from aviation-related UFP reductions.

Aviation-related air pollution

Generalized fiducial inference on differentiable manifolds

We introduce a novel approach to inference on parameters that take values in a Riemannian manifold embedded in a Euclidean space. Parameter spaces of this form are ubiquitous across many fields, including chemistry, physics, computer graphics, and geology. Here, this new approach uses generalized fiducial inference (GFI) to obtain a posterior-like distribution on the manifold, without needing to know local parameterizations that map to the constrained space from an unconstrained Euclidean space. Using mathematical tools from Riemannian geometry, we construct a constrained generalized fiducial distribution (CGFD). A Bernstein-von Mises-type result for the CGFD, which provides intuition for how the desirable asymptotic qualities of the unconstrained generalized fiducial distribution are inherited by the CGFD, is provided. To illustrate the practical use of the CGFD, we provide a proof-of-concept example in the context of a linear logspline density estimation problem, and demonstrate that CGFD-based confidence sets exhibit desirable coverage properties via simulation. As an application, we fit a CGFD to COVID-19 case count data from North Carolina, USA.

97 MATHEMATICS AND COMPUTING

Prime VI

SAND2025-03757O Prime VI is a distribution-of-disease outbreak model calibration code based on variational inference. It accompanies a publication for submission to Statistics in Medicine journal, and the code will be maintained for open-source use on Sandia's GitLab. The software provides methods for calibrating an epidemiological model to measured case-count data for a multitude of correlated spatial regions. The code solves a Bayesian inverse problem for model calibration where the posterior over-model parameters are approximated through a custom implementation of variational inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING

Near-Efficient and Non-Asymptotic Multiway Inference

We establish non-asymptotic efficiency guarantees for tensor decomposition–based inference in count data models. Under a Poisson framework, we consider two related goals: (i) parametric inference , the estimation of the full distributional parameter tensor, and (ii) multiway analysis , the recovery of its canonical polyadic (CP) decomposition factors. Our main result shows that in the rank-one setting, a rank-constrained maximum-likelihood estimator achieves multiway analysis with variance matching the Cramér–Rao Lower Bound (CRLB) up to absolute constants and logarithmic factors. This provides a general framework for studying “near-efficient” multiway estimators in finite-sample settings. For higher ranks, we illustrate that our multiway estimator may not attain the CRLB; nevertheless, CP-based parametric inference remains nearly minimax optimal, with error bounds that improve on prior work by offering more favorable dependence on the CP rank. Numerical experiments corroborate near-efficiency in the rank-one case and highlight the efficiency gap in higher-rank scenarios.

97 MATHEMATICS AND COMPUTING

The Poisson tensor completion parametric estimator

We introduce the Poisson tensor completion (PTC) estimator that exploits inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram for samples of a multivariate distribution. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial non-homogeneous Poisson process. The Poisson tensor decomposition leads to a completion of the mean measure over all bins—including those containing few to no samples—and leads to our proposed estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values obviating the need for additional constraints to ensure non-negativity. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

97 MATHEMATICS AND COMPUTING

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195

SIMS Data Correction Procedure for Quasi‐Simultaneous Arrival ( QSA ) Under‐counting and Ramifications of Misapplication

Isotope geochemistry requires isotope ratios measured using secondary ion mass spectrometry (SIMS) to be made with optimal precision and accuracy. Under some analytical conditions when using electron multiplier detectors, secondary ions may be under‐counted because of quasi‐simultaneous arrival (QSA) at the first dynode. The relative magnitude of the associated QSA correction to raw measured isotopic ratios can be up to seventy permil or more. Therefore, not applying the correction, or misapplication of it could lead to significant inaccuracies in published isotope ratio data. Examples and ramifications of the latter are described in addition to a straightforward procedure for QSA under‐counting correction.

Geochemistry & Geophysics

Poisson-response Tensor-on-Tensor Regression and Applications

We introduce Poisson-response tensor-on-tensor regression (PToTR), a novel regression framework designed to handle tensor responses composed element-wise of random Poisson-distributed counts. Tensors, or multi-dimensional arrays, composed of counts are common data in fields such as inter national relations, social networks, epidemiology, and medical imaging, where events occur across multiple dimensions like time, location, and dyads. PToTR accommodates such tensor responses alongside tensor covariates, providing a versatile tool for multi dimensional data analysis. We propose algorithms for maximum likelihood estimation under a canonical polyadic (CP) structure on the regression coefficient tensor that satisfy the positivity of Poisson parameters and then provide an initial theoretical error analysis for PToTR estimators. We also demonstrate the utility of PToTR through three concrete applications: longitudinal data analysis of the Integrated Crisis Early Warning System database, positron emission tomography (PET) image reconstruction, and change-point detection of communication patterns in longitudinal dyadic data. These applications highlight the versatility of PToTR in addressing complex, structured count data across various domains.

97 MATHEMATICS AND COMPUTING

Soil viral production count, respiration, and amplicon data

This study aimed to quantify rates of viral production in aridisol soil under conditions as close to natural field soil as possible given the perturbations necessary to manipulate viral abundances. Viruses were removed from soil, then added back to virus-depleted soil to control the initial viral abundances at either 100% (field_abund) or 10% (reduced_abund) of measured field viral abundance to generate treatments with field-relevant and reduced viral infection pressure, respectively. Replicates of batch incubation jars were harvested every 8 hours for 48 hours to enumerate bacteria and viruses by microscopy (n=5) and profile bacterial community composition by 16S rRNA amplicon sequencing (n=3).

Zimmerman, Amy [Pacific Northwest National Laborat