Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “applied statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Toward In Situ Synchrotron Mapping of Crystal Selection Processes during Crystal Growth

In this work, we present an automated and rapid method for non-destructive mapping of crystal grains in a rod-shaped sample. The approach was designed for application to in situ float-zone crystal growth experiments at an x-ray synchrotron source, but could be useful in other applications. The methods have been tested on a TiO 2 boule grown in an optical float zone furnace. The approach applies a statistical filter to polycrystalline diffraction patterns on 2D detectors to rapidly determine the degree of powder-quality 1of the signal. When larger crystals emerge in the growth, their position, size and shape can be tracked using an automated blob-tracking algorithm that follows individual Bragg peaks as a function of position in a grid-scan, even when multiple crystals are contributing spots to diffraction images. This method is found to be robust as the same crystal shape can be independently reconstructed using different sets of Bragg reflections. Image segmentation methods are then used to map out the polycrystalline grains. We also note that other information about crystal quality, such as mosaicity or strain state, may be inferred and mapped from the intensity variation of the Bragg peaks at different locations within the sample.

36 MATERIALS SCIENCE↗

Size-dependent shape distributions of platinum nanoparticles

Transmission electron microscopy revealed size-dependent shape distributions in platinum nanoparticles, which were consistent with trends observed by applying Boltzmann statistics to the energy computed with atomistic models.

36 MATERIALS SCIENCE↗

Redshifting galaxies from DESI to JWST CEERS: Correction of biases and uncertainties in quantifying morphology

Observations of high-redshift galaxies with unprecedented detail have now been rendered possible with the James Webb Space Telescope (JWST). However, accurately quantifying their morphology remains uncertain due to potential biases and uncertainties. To address this issue, we used a sample of 1816 nearby DESI galaxies, with a stellar mass range of 10 9.75 - 11.25 M ⊙ , to compute artificial images of galaxies of the same mass located at 0.75 ≤ z ≤ 3 and observed at rest-frame optical wavelength in the Cosmic Evolution Early Release Science (CEERS) survey. We analyzed the effects of cosmological redshift on the measurements of Petrosian radius (R p ), half-light radius (R 50 ), asymmetry (A), concentration (C), axis ratio (q), and Sérsic index (n). Our results show that R p and R 50 , calculated using non-parametric methods, are slightly overestimated due to PSF smoothing, while R 50 , q, and n obtained through fitting a Sérsic model does not exhibit significant biases. By incorporating a more accurate noise effect removal procedure, we improve the computation of A over existing methods, which often overestimate, underestimate, or lead to significant scatter of noise contributions. Due to PSF asymmetry, there is a minor overestimation of A for intrinsically symmetric galaxies. However, for intrinsically asymmetric galaxies, PSF smoothing dominates and results in an underestimation of A, an effect that becomes more significant with higher intrinsic A or at lower resolutions. Moreover, PSF smoothing also leads to an underestimation of C, which is notably more pronounced in galaxies with higher intrinsic C or at lower resolutions. We developed functions based on resolution level, defined as R p /FWHM, for correcting these biases and the associated statistical uncertainties. Applying these corrections, we measured the bias-corrected morphology for the simulated CEERS images and we find that the derived quantities are in good agreement with their intrinsic values – except for A, which is robust only for angularly large galaxies where R p /FWHM ≥ 5. Our correction functions can be applied to other surveys, offering valuable tools for future studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Decoding the age–chemical structure of the Milky Way disc: an application of copulas and elicitable maps

In the Milky Way, the distribution of stars in the [α/Fe] versus [Fe/H] and [Fe/H] versus age planes holds essential information about the history of star formation, accretion, and dynamical evolution of the Galactic disc. We investigate these planes by applying novel statistical methods called copulas and elicitable maps to the ages and abundances of red giants in the Apache Point Observatory Galactic Evolution Experiment survey. We find that the high- and low-α disc stars have a clean separation in copula space and use this to provide an automated separation of the α sequences using a purely statistical approach. This separation reveals that the high-α disc ends at the same [α/Fe] and age at high [Fe/H] as the low-[Fe/H] start of the low-α disc, thus supporting a sequential formation scenario for the high- and low-α discs. We then combine copulas with elicitable maps to precisely obtain the correlation between stellar age τ and metallicity [Fe/H] conditional on Galactocentric radius R and height z in the range 0 < R < 20 kpc and |z| < 2 kpc. The resulting trends in the age–metallicity correlation with radius, height, and [α/Fe] demonstrate a ≈0 correlation wherever kinematically cold orbits dominate, while the naively expected negative correlation is present where kinematically hot orbits dominate. This is consistent with the effects of spiral-driven radial migration, which must be strong enough to completely flatten the age–metallicity structure of the low-α disc.

79 ASTRONOMY AND ASTROPHYSICS↗

Multiplicative Shot-Noise: A New Route to Stability of Plastic Networks

Fluctuations of synaptic weights, among many other physical, biological, and ecological quantities, are driven by coincident events of two “parent” processes. Here we propose a multiplicative shot-noise model that can capture the behaviors of a broad range of such natural phenomena, and analytically derive an approximation that accurately predicts its statistics. We apply our results to study the effects of a multiplicative synaptic plasticity rule that was recently extracted from measurements in physiological conditions. Using mean-field theory analysis and network simulations, we investigate how this rule shapes the connectivity and dynamics of recurrent spiking neural networks. The multiplicative plasticity rule is shown to support efficient learning of input stimuli, and it gives a stable, unimodal synaptic-weight distribution with a large fraction of strong synapses. The strong synapses remain stable over long times but do not “run away.” Our results suggest that the multiplicative shot-noise offers a new route to understand the tradeoff between flexibility and stability in neural circuits and other dynamic networks.

59 BASIC BIOLOGICAL SCIENCES↗

Towards Quantum Computing Phase Diagrams of Gauge Theories with Thermal Pure Quantum States

The phase diagram of strong interactions in nature at finite temperature and chemical potential remains largely theoretically unexplored due to inadequacy of Monte-Carlo–based computational techniques in overcoming a sign problem. Quantum computing offers a sign-problem-free approach, but evaluating thermal expectation values is generally resource intensive on quantum computers. To facilitate thermodynamic studies of gauge theories, we propose a generalization of the thermal-pure-quantum-state formulation of statistical mechanics applied to constrained gauge-theory dynamics, and numerically demonstrate that the phase diagram of a simple low-dimensional gauge theory is robustly determined using this approach, including mapping a chiral phase transition in the model at finite temperature and chemical potential. Quantum algorithms, resource requirements, and algorithmic and hardware error analysis are further discussed to motivate future implementations. Thermal pure quantum states, therefore, may present a suitable candidate for efficient thermal simulations of gauge theories in the era of quantum computing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

EFBMC

SAND2026-23986O EFBMC performs elastic Bayesian model calibration by applying Bayesian statistics and functional analysis. The software provides Python, R, and MATLAB scripts that enable users to calibrate models. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗

High-Burnup LOCA Burst Susceptibility BISON Analysis in PWRs and BWRs

Accurately assessing high-burnup fuel behavior during loss-of-coolant accidents (LOCAs) is essential for understanding fuel fragmentation, relocation, and dispersal (FFRD) risks across the US light-water reactor fleet. This work updates previous Nuclear Energy Advanced Modeling and Simulation (NEAMS) Program multiphysics LOCA analyses for a pressurized water reactor (PWR) and a boiling water reactor (BWR) by incorporating recent model and material property advancements in the BISON fuel performance code, including a high-burnup structure (HBS) model, revised cladding burst criteria, and updated thermal–mechanical correlations. This update was needed to support ongoing industry initiatives and upcoming regulatory changes. Full-core, rod-resolved operating histories generated using Virtual Environment for Reactor Analysis (VERA) and system-level LOCA conditions obtained from TRACE were applied to statistically representative rod samples in BISON to evaluate burst behavior and FFRD susceptibility. These calculations used two cladding burst correlations and three fuel pulverization models so that the predictions of these models could be compared. The updated PWR simulations show markedly improved numerical stability as the number of crashed simulations decreased by 95% compared to the previous study, and hence higher confidence in results. The updated PWR simulations predicted cladding bursts exclusively among once-burned, high-power rods, with two different cladding burst models identifying the same burst-susceptible population. Resulting FFRD susceptibility estimates are significantly reduced compared with earlier studies, driven by cooler predicted fuel and plenum temperatures, lower hoop strains, and reduced fission gas release in the updated models. In contrast, none of the BWR rods were predicted to burst under either burst criterion, reaffirming minimal BWR FFRD susceptibility even with updated HBS and material models. Comparisons between the PWR and BWR end-of-cycle predictions are made. Comparison with prior work highlights significant shifts in PWR fuel performance metrics and confirmation of earlier BWR conclusions. Overall, the updated results underscore the importance of having high-resolution detailed modeling capability and continuously integrating evolving material models and physics into high-resolution multiphysics simulations. The unified assessment presented here strengthens confidence in predicting high-burnup LOCA behavior by improving agreement between different cladding burst correlations. These results also provide an improved foundation for future BISON model development, FFRD susceptibility calculations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Catalog of Candidate Double and Lensed Quasars from Gaia and WISE Data

Making use of strong correlations between closely separated multiple or double sources and photometric and astrometric metadata in Gaia Early Data Release 3 (EDR3), we generate a catalog of candidate double- and multiply imaged lensed quasars and active galactic nuclei (AGNs), comprising 3140 systems. It includes two partially overlapping parts: a sample of distant (redshifts mostly greater than 1) sources with perturbed data; and systems that have been resolved into separate components by Gaia at separations less than 2''. For the first part, which is roughly one-third of the published catalog, we synthesized 0.617 million redshifts using multiple machine-learning prediction and classification methods, using independent photometric and astrometric data from Gaia EDR3 and the Wide-field Infrared Survey Explorer, with accurate spectroscopic redshifts from the Sloan Digital Sky Survey (SDSS) as a training set. Using these synthetic redshifts, we estimate a 4.9% rate of interlopers with spectroscopic redshifts below 1 in this part of the catalog. Unresolved candidate double and dual AGNs and quasars are selected as sources with a marginally high BP/RP excess factor (phot_bp_rp_excess_factor), which is sensitive to source extent, limiting our search to high-redshift quasars. For the second part of the catalog, additional filters on measured parallax and near-neighbor statistics are applied to diminish the propagation of the remaining stellar contaminants. The estimated rate of the positives (double or multiple sources) is 98%, and the estimated rate of dual (physically related) quasars is greater than 54%. A few dozen serendipitously found objects of interest are discussed in more detail, including known and new lensed images, planetary nebulae, young IR stars of peculiar morphology, and quasars with catastrophic redshift errors in SDSS.

79 ASTRONOMY AND ASTROPHYSICS↗

Flexible and Adaptive Malware Identification Using Techniques from Biology

The holy grail in cyber analytics is to find new ways to understand the information we already have access to. One way to do that is to characterize the data into reasonable sizes and then leverage any known information to generate new insights. Biologists have been using a similar process for decades. This paper introduces the MLSTONES tool set that was developed by leveraging biology and bioinformatics, high performance computing, and statistical algorithms applied to cyber data and specifically to malware. Furthermore, the paper discusses the tool suite, its applications, and how it compares or can work with other related tools.

Peterson, Elena S.↗

IDENTIFICATION OF POTENTIAL SUPERCONDUCTOR QUENCH PRECURSORS USING FREQUENCY DOMAIN FEATURE ANALYSIS

Superconducting magnets are important pieces of technology in the world of particle accelerators, allowing researchers to study atomic and subatomic phenomena, among other things. In some instances, superconductors can lose this non-resistive property in a phenomenon known as quenching, which can cause damage to the magnets. This potential danger prompts the introduction of systems to predict when a quench is imminent; one such implementation is through the use of acoustic sensors that detect vibrations within the magnet. Within these acoustic sensor signals, significantly above-noise disturbances (referred to as ”events”) can be identified. Our research applies the statistical framework of a permutation test to features calculated from the power spectral density (PSD) to distinguish between events far from the quench at the end of the signal to events at the start of the signal. We found that dividing the PSD into frequency bands produced a feature capable of distinguishing between events early in the signal and late in the signal leading up the quench, providing a promising starting place for future quench prediction systems.

Roehrig, Benjamin [Northern Illinois U.]↗

2D k -th nearest neighbour statistics: a highly informative probe of galaxy clustering

ABSTRACT Beyond standard summary statistics are necessary to summarize the rich information on non-linear scales in the era of precision galaxy clustering measurements. For the first time, we introduce the 2D k-th nearest neighbour (kNN) statistics as a summary statistic for discrete galaxy fields. This is a direct generalization of the standard 1D kNN by disentangling the projected galaxy distribution from the redshift-space distortion signature along the line-of-sight. We further introduce two different flavours of 2D kNNs that trace different aspects of the galaxy field: the standard flavour which tabulates the distances between galaxies and random query points, and a ‘DD’ flavour that tabulates the distances between galaxies and galaxies. We showcase the 2D kNNs’ strong constraining power both through theoretical arguments and by testing on realistic galaxy mocks. Theoretically, we show that 2D kNNs are computationally efficient and directly generate other statistics such as the popular two-point correlation function (2PCF), voids probability function, and counts-in-cell statistics. In a more practical test, we apply the 2D kNN statistics to simulated galaxy mocks that fold in a large range of observational realism and recover parameters of the underlying extended halo occupation distribution (HOD) model that includes velocity bias and galaxy assembly bias. We find unbiased and significantly tighter constraints on all aspects of the HOD model with the 2D kNNs, both compared to the standard 1D kNN, and the classical redshift-space 2PCF.

79 ASTRONOMY AND ASTROPHYSICS↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bayesian estimation of the $S$ factor and thermonuclear reaction rate for 16 O(p, γ) 17 F

The 16 O(p, γ) 17 F reaction is the slowest hydrogen-burning process in the CNO mass region. Its thermonuclear rate sensitively impacts predictions of oxygen isotopic ratios in a number of astrophysical sites, including AGB stars. The reaction has been measured several times at low bombarding energies using a variety of techniques. The most recent evaluated experimental rates have a reported uncertainty of about 7.5% below 1 GK. However, the previous rate estimate represents a best guess only and was not based on rigorous statistical methods. We apply a Bayesian model to fit all reliable 16 O(p, γ) 17 F cross section data, and take into account independent contributions of statistical and systematic uncertainties. The nuclear reaction model employed is a single-particle potential model involving a Woods-Saxon potential for generating the radial bound state wave function. The model has three physical parameters, the radius and diffuseness of the Woods-Saxon potential, and the asymptotic normalization coefficients (ANCs) of the final bound state in 17 F. Here, we find that performing the Bayesian S -factor fit using ANCs as scaling parameters has a distinct advantage over adopting spectroscopic factors instead. Based on these results, we present the first statistically rigorous estimation of experimental 16 O(p, γ) 17 F reaction rates, with uncertainties (±4.2%) of about half the previously reported values.

6 ≤ A ≤ 19↗

KAZRARSCL-CLOUDSAT Value Added Product Data Stream

The KAZRARSCL-CLOUDSAT Value-Added Product (VAP) is based on the KAZR-ARSCL VAP, which provides cloud boundaries and best-estimate time-height fields of radar moments. The KAZRARSCL-CLOUDSAT VAP applies a statistically-derived calibration offset to reflectivity fields in order to align them with observations from the spaceborne CloudSat Cloud Profiling Radar. For details on the offset values applied, please refer to "Kollias, P., Puigdomènech Treserras, B., and Protat, A.: Calibration of the 2007–2017 record of ARM Cloud Radar Observations using CloudSat, Atmos. Meas. Tech. Discuss., https://doi.org/10.5194/amt-2019-34, 2019."

54 ENVIRONMENTAL SCIENCES↗

How Climate and Data Quality Impact Photovoltaic Performance Loss Rate Estimations

Different data pipelines and statistical methods are applied to photovoltaic (PV) performance datasets to quantify the performance loss rate (PLR). Since the real values of PLR are unknown, a variety of unvalidated values are reported. As such, the PV industry commonly assumes PLR based on statistically extracted ranges from the literature. However, the accuracy and uncertainty of PLR depend on several parameters including seasonality, local climatic conditions, and the response of a particular PV technology. In addition, the specific data pipeline and statistical method used affect the accuracy and uncertainty. To provide insights, a framework of (≈200 million) synthetic simulations of PV performance datasets using data from different climates is developed. Time series with known PLR and data quality are synthesized, and large parametric studies are conducted to examine the accuracy and uncertainty of different statistical approaches over the contiguous US, with an emphasis on the publicly available and “standardized” library, RdTools . In the results, it is confirmed that PLRs from RdTools are unbiased on average, but the accuracy and uncertainty of individual PLR estimates vary with climate zone, data quality, PV technology, and choice of analysis workflow. Best practices and improvement recommendations based on the findings of this study are provided.

14 SOLAR ENERGY↗

Mathematical nuances of Gaussian process-driven autonomous experimentation

Abstract The fields of machine learning (ML) and artificial intelligence (AI) have transformed almost every aspect of science and engineering. The excitement for AI/ML methods is in large part due to their perceived novelty, as compared to traditional methods of statistics, computation, and applied mathematics. But clearly, all methods in ML have their foundations in mathematical theories, such as function approximation, uncertainty quantification, and function optimization. Autonomous experimentation is no exception; it is often formulated as a chain of off-the-shelf tools, organized in a closed loop, without emphasis on the intricacies of each algorithm involved. The uncomfortable truth is that the success of any ML endeavor, and this includes autonomous experimentation, strongly depends on the sophistication of the underlying mathematical methods and software that have to allow for enough flexibility to consider functions that are in agreement with particular physical theories. We have observed that standard off-the-shelf tools, used by many in the applied ML community, often hide the underlying complexities and therefore perform poorly. In this paper, we want to give a perspective on the intricate connections between mathematics and ML, with a focus on Gaussian process-driven autonomous experimentation. Although the Gaussian process is a powerful mathematical concept, it has to be implemented and customized correctly for optimal performance. We present several simple toy problems to explore these nuances and highlight the importance of mathematical and statistical rigor in autonomous experimentation and ML. One key takeaway is that ML is not, as many had hoped, a set of agnostic plug-and-play solvers for everyday scientific problems, but instead needs expertise and mastery to be applied successfully. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

Hydrogen Dispersion Modeling for Development of Smart Distributed Monitoring

Studying hydrogen dispersion is crucial for ensuring the safe and effective deployment of hydrogen as an energy carrier. This study presents a comprehensive CFD modeling framework for simulating hydrogen dispersion at a real-world hydrogen production, storage, and utilization facility. Utilizing the Hydrogen Research Facility under the Advanced Research on Integrated Energy Systems (ARIES) at the National Renewable Energy Laboratory's (NREL) Flatirons campus, controlled hydrogen releases at 27 kg-H2/hr were simulated. The model incorporated site-specific atmospheric conditions, including hourly wind speeds and temperatures recorded between 8 AM and 8 PM from October to December 2023. To reduce computational demands, a statistical reduction technique was applied to condense the dataset to 100 representative scenarios, validated by statistical tests for wind speeds and power law coefficients. Simulations were conducted using the Reynolds-Averaged Navier-Stokes equations. Results demonstrated that wind speed substantially influences hydrogen dispersion, with low wind conditions forming concentrated clouds and higher wind speeds stretching the plume. Additionally, clustering analysis informed optimal sensor placement at various elevations with up to 10 sensor locations on each elevation. This framework offers a robust approach for understanding hydrogen behavior in ambient conditions and informing detection strategies.

08 HYDROGEN↗