Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “histograms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.

97 MATHEMATICS AND COMPUTING↗

Direct current response of a thin scCVD diamond detector under increased applied field to 14.1 MeV neutrons

An avalanche effect yielding inherent gain can be exploited in thin, single-crystal chemical vapor deposition (scCVD) diamond. It occurs when a high enough bias is applied across the diamond thickness while avoiding breakdown. This charge multiplication effect was studied previously with alpha particles and heavy ions either by using the transient current technique or by measuring the energy spectrum. The measurements we obtained to evaluate the charge multiplication performance of a 10 μm thick scCVD diamond detector used a novel approach—we employed an electrometer to characterize the response of the detector by performing directly coupled current measurements (time-averaged charge, at 1 Hz sampling) when exposed to 14.1 MeV neutrons from deuterium-tritium fusion. We measured both the dark and irradiated currents from the detector over a range of applied displacement field values from 2 to 75 V/μm. A histogram method with central mean and standard deviation width was used to determine the current over each measurement duration typically from 100 to 300 seconds. The dark-subtracted irradiated current (i.e., contrast) was used to evaluate the gain of the detector at each applied displacement field. The contrast at an applied displacement field between 15 and 20 V/μm was higher than the expected linear increase in contrast proportional to the increased applied bias, indicating the possible presence of avalanche events in the diamond. The detector response also indicated possible polarization and charge depletion effects. These results provide an opportunity to further explore the use of thin scCVD diamond as a fast neutron current mode detector with inherent gain.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A practical guide to unbinned unfolding

Unfolding, in the context of high-energy particle physics, refers to the process of removing detector distortions in experimental data. The resulting unfolded measurements are straightforward to use for direct comparisons between experiments and a wide variety of theoretical predictions. For decades, popular unfolding strategies were designed to operate on data formatted as one or more binned histograms. In recent years, new strategies have emerged that use machine learning to unfold datasets in an unbinned manner, allowing for higher-dimensional analyses and more flexibility for current and future users of the unfolded data. This guide comprises recommendations and practical considerations from researchers across a number of major particle physics experiments who have recently put these techniques into practice on real data.

Canelli, Florencia [Univ. of Zurich (Switzerland)]↗

Poster Abstract: Leveraging Large Language Models to Reveal Interpretable Cooling Behaviors from Smart Thermostat Data

Frequent heatwaves and hot summers increasingly challenge occupant comfort, health, and energy grid stability. Addressing these challenges requires a detailed understanding of household cooling behaviors, such as thermostat adjustments and adaptive responses to extreme conditions. Traditional analyses often rely on aggregated numerical metrics that overlook subtle but important household-specific variations. In this study, we introduce a generalizable methodology that integrates large language models (LLMs) with vision capabilities to enable scalable and detailed analysis of residential thermostat data. Using Ecobee's Donate Your Data (DYD) dataset—which provides five-minute records of indoor temperatures, thermostat setpoints, and HVAC runtimes—we focus on two U.S. cities with contrasting summer climates : Austin (TX) and Phoenix (AZ). Because raw time-series data are not well suited for direct LLM analysis, we transform them into visual representations, such as daily indoor temperature trajectories and weekly runtime histograms, to better capture behavioral variations. Leveraging LLMs' visual interpretation, we extract descriptive behavioral features, including temperature preferences, time-of-day cooling orientation, anticipatory versus reactive heatwave responses, and behavioral consistency. These semantic features support unsupervised clustering to identify distinct occupant archetypes at scale, revealing differences—such as morning-centric anticipatory coolers versus households that shift toward warmer setpoints during heatwaves—that can inform demand response, resilience planning, and health-aware interventions. By converting raw numerical data into interpretable behavioral patterns, this methodology enables scalable and practical analysis of occupant behavior, supporting actionable insights for comfort, resilience, and energy management.

Nihar, Kopal↗

Observation of a single protein by ultrafast X-ray diffraction

* runs/ - The raw data from the pnCCD * histograms.h5 - Histogram of the values recorded by each pixel throughout the data acquisition. Used for pedestal and gain correction. * mask.h5 - Detector bad pixels mask. * 1SS8.pdb - PDB model of the assembled GroEL 14-mer, from the 1SS8 PDB entry. * 1SS8_H2O.pdb - Hydrated version of the assembled GroEL 14-mer, derived from the 1SS8 PDB entry. The various hydrated models are created by removing water from this initial model. * rotations.h5 - Quaternions of the rotations used for template matching

European XFEL↗

XRD-PUAT (X-Ray Diffraction - Parameter Uncertainty Analysis Toolkit)

This XRD uncertainty toolkit was built to investigate uncertainty and local least squares topology over a specified parameter space on a single histogram in a gpx GSAS-II file. The parameter uncertainty can be investigated using either a frequentist F-test approach or a Bayesian Inference statistical inversion method on a weighted least squares or peak fit refinements.SAND Number: SAND2020-12227 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Moore, Alexander↗

HUNTRESS: a fast heuristic for reconstructing phylogenetic trees of tumor evolution (HUNTRESS) v0.1

We introduce HUNTRESS (Histogrammed UNion Tree REconStruction heuriStic), a computational method for tumor phylogeny reconstruction from noisy genotype matrices derived from single-cell sequencing data, whose running time is linear with the number of cells and quadratic with the number of mutations. Provided that the input genotype matrix includes no false positives, each cellular subpopulation is at least a user defined fraction of the total number of cells, and the number of cells are much bigger than the number of mutations considered, HUNTRESS computes the ground truth tumor phylogeny with high probability. On simulated data HUNTRESS is faster than available alternatives with comparable or better accuracy. Additionally, the phylogenies reconstructed by HUNTRESS on two single-cell sequencing data sets agree with the best known evolutionary scenarios for the associated tumors.

Buluc, Aydin↗

Tidal Disruption Event Galaxy Binner

This software simulates astronomical survey detections of tidal disruptions of stars by super-massive black holes. It begins with the synthetic galaxy catalogue described in van Velzen 2008 (https://arxiv.org/abs/1707.03458). The stellar disruption rate in each galaxy is estimated based on Stone & Metzger 2016 (https://arxiv.org/abs/1410.7772). Based on these rates, and the present-day stellar mass function in the galaxy, disruptions are randomly sampled, and the properties of the resulting flares are sampled based on empirical distributions. The code also accounts for obscuration by dust in the host galaxy. Finally, the survey selection effects are applied. The detectable simulated flares are stored in a database, allowing histograms of their properties to be created.

Roth, NathanielJ.↗

PRESSURE-FILM OPEN SOURCE SCANNING AND MAPPING

SF-22-099 Pressure-film Open Source Scanning and Mapping (POSSM) is software that analyses images of pressure-film (scans or photographs). POSSM collects data from that analysis in order to produce a number of visual aides, including pressure maps, histograms, and 3d surface plots. POSSM is designed to be able to analyze an entire folder full of images, exporting numeric data as .csv files and visual graphics as .png files.

Grider, Patrick↗

Simulation-Based Inference for Neutrino Interaction Model Tuning

This project demonstrates, for the first time, the application of simulation-based inference (SBI) techniques to tune neutrino–nucleus interaction models. Using a mock dataset based on the MicroBooNE tuning of the GENIE event generator, our approach employs a Neural Posterior Estimator (NPE) with Masked Autoregressive Flows (MAF) to infer the posterior distributions of key GENIE parameters directly from simulated histograms. The workflow provides a scalable and amortized framework for performing likelihood-free inference in high-dimensional parameter spaces, offering a pathway to more efficient and uncertainty-aware model tuning for next-generation neutrino experiments such as DUNE and SBND.

Tame-Narvaez, KarlaMaria [Fermi National Accelerat↗

Simulations of ENSO Phase-locking in CMIP5 and CMIP6

The characteristics of El-Niño-Southern Oscillation (ENSO) phase-locking in observations and CMIP5 and CMIP6 models are examined in this study. Two metrics based on the peaking month histogram for all El Niño and La Niña events are adopted to delineate the basic features of ENSO phase-locking in terms of the preferred calendar month and strength of this preference. It turns out that most models are poor at simulating the ENSO phase-locking, either showing little peak strengths or peaking at the wrong seasons. By deriving ENSO’s linear dynamics based on the conceptual recharge oscillator (RO) framework through the seasonal linear inverse model (sLIM) approach, various simulated phase-locking behaviors of CMIP models are systematically investigated in comparison with observations. In observations, phase-locking is mainly attributed to the seasonal modulation of ENSO’s SST growth rate. In contrast, in a significant portion of CMIP models, phase-locking is co-determined by the seasonal modulations of both SST growth and phase-transition rates. Further study of the joint effects of SST growth and phase-transition rates suggests that for simulating realistic winter peak ENSO phase-locking with the right dynamics, climate models need to have four key factors in the right combination: (1) correct phase of SST growth rate modulation peaking at the fall; (2) large enough amplitude for the annual cycle in growth rate; (3) amplitude of semi-annual cycle in growth rate needs to be small; and (4) amplitude of seasonal modulation in SST phase-transition rate needs to be small.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Two Large-Scale Meteorological Patterns are Associated with Short-Duration Dry Spells in the Northeastern United States

Large-scale meteorological pattern (LSMP)–based analysis is used in a novel way to understand meteorological conditions before and during short-duration dry spells over the northeastern United States. These LSMPs are useful to assess models and select better-performing models for future projections. Dry-spell events are identified from histograms of consecutive dry days below a daily precipitation threshold. Events lasting 12 days or longer, which correspond to ~10% of dry-spell events, are examined. The 500-hPa stream-function anomaly fields for the first 12 days of each event are time averaged, and k -means clustering is applied to isolate the dry-spell-related LSMPs. The first cluster has a strong low pressure anomaly over the Atlantic Ocean, southeast of the region, and is more common in winter and spring. The second cluster has strong high pressure over east-central North America and is most common during autumn. Over the region, both clusters have negative specific humidity anomalies, negative integrated vapor transport from the north, and subsidence associated with a midlatitude jet stream dipole structure that reinforces upper-level convergence. Subsidence is supported by cold-air advection in the first cluster and the location on the east side of the lower-level high pressure in the second cluster. Extratropical cyclone storm tracks are generally shifted southward of the region during the dry spells. Individual events lie on a continuum between two distinct clusters. These clusters have similar local, but different remote, properties. Although dry spells occur with greater frequency during drought months, most dry spells occur during nondrought months. Significance Statement: This study examines the large-scale weather patterns and meteorological conditions associated with dry-spell events lasting at least 2 weeks while affecting the northeastern United States. A statistical approach groups events together on the basis of similar atmospheric features. We find two distinct sets of patterns that we call large-scale meteorological patterns. These patterns reduce moisture, foster localized sinking, and shift the storm track southward along the Atlantic seaboard, all of which reduce precipitation. Besides greater understanding, knowing the meteorological patterns during short-term dryness in the region provides an important tool to assess how well atmospheric models reproduce these specific patterns. More dry spells occur in non-drought months than in drought months, which means that dry spells can occur without preexisting drought conditions.

58 GEOSCIENCES↗

DEEPEN 3D PFA Index Models for Exploration Datasets at Newberry Volcano

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), index models needed to be developed to map values in geoscientific exploration datasets to favorability index values. This GDR submission includes those index models. Index models were created by binning values in exploration datasets into chunks based on their favorability, and then applying a number between 0 and 5 to each chunk, where 0 represents very unfavorable data values and 5 represents very favorable data values. To account for differences in how exploration methods are used to detect each play component, separate index models are produced for each exploration method for each component of each play type. Index models were created using histograms of the distributions of each exploration dataset in combination with literature and input from experts about what combinations of geophysical, geological, and geochemical signatures are considered favorable at Newberry. This is in attempt to create similar sized bins based on the current understanding of how different anomalies map to favorable areas for the different types of geothermal plays (i.e., conventional hydrothermal, superhot EGS, and supercritical). For example, an area of partial melt would likely appear as an area of low density, high conductivity, low vp, and high vp/vs. This means that these target anomalies would be given high (4 or 5) index values for the purpose of imaging the heat source. To account for differences in how exploration methods are used to detect each play component, separate index models are produced for each exploration method for each component of each play type. Index models were produced for the following datasets: - Geologic model - Alteration model - vp/vs - vp - vs - Temperature model - Seismicity (density*magnitude) - Density - Resistivity - Fault distance - Earthquake cutoff depth model

15 GEOTHERMAL ENERGY↗

NOvA 2020 official data release (13.6E20 neutrino + 12.5E20 antineutrino)

NOvA results of the 2020 joint analysis of nue appearance and numu disappearance data presented in Phys. Rev. D 106, 032004:https://doi.org/10.1103/PhysRevD.106.032004Note: these contained plots are slightly different from the ones shown at NEUTRINO 2020.Exposure: 13.6E20 protons on target, neutrino-enhanced beam12.5E20 protons on target, antineutrino-enhanced beamParameters not constrained by NOvA:- th13 is assumed to be Gaussian, ss2th13 = 0.085 +/- 0.003 (PDG 2019, Phys. Rev. D98, 030001 (2018) and 2019 update, https://pdg.lbl.gov/2019/listings/rpp2019-list-neutrino-mixing.pdf)- ss2th12 = 0.851, dm21=7.53e-5 are fixed at 2019 PDG valuesSignificances are obtained with the Feldman-Cousins procedure.Files contain TGraphs for the following:- ssth23 vs deltaCP 1,2,3 sigma and 90% CL contours, for NO or IO- dmsq32 vs ssth23 1,2,3 sigma and 90% CL contours, for NO or IO- deltaCP profile for NO UO, NO LO, IO LO, IO UO- ssth23 profile for NO, IO- dmsq32 profile for NO UO, NO LO, IO LO, IO UOContour files also include a TMarker with the best-fit parameters (in the NO)An additional file containing the predictions for all channels’ signal and background components computed at the NOvA best-fit oscillation parameters and systematic pull terms. Numu samples include the no-oscillation case as well.Please note that any fits performed with these histograms are not expected to exactly reproduce the official NOvA results as parameterizations of the numerous systematic uncertainties considered in the official fits are not included in this release.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spatial Study 2022: Water Column, Sediment, and Total Ecosystem Respiration Rates across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin and is associated with the manuscript “Sediment-associated processes account for most of the spatial variation in ecosystem respiration in the Yakima River basin” submitted to Nature Communications Earth & Environment (Garayburu-Caruso et al., in review). The dataset provides ecosystem metabolism estimates generated from streamMetabolizer (Appling et al.; 2018) using data collected during the same five-week period at 48 sites within multiple rivers throughout the Yakima River Basin in Washington, USA. Additionally, it includes the scripts used for the analysis and producing the figures in the manuscript. The contents include streamMetabolizer inputs and outputs and additional relevant data needed to generate the main manuscript results. The data included are: total ecosystem respiration, water respiration, calculated sediment-associated respiration, gross primary production outputs from the river corridor model for the Yakima River Basin, median grain size (d50), depth, dissolved oxygen, water temperature, pressure, and annual oxygen consumption. The associated GitHub repository can be found at https://github.com/river-corridors-sfa/SSS_metabolism. Samples collected during this study were labeled as “Second Spatial Study” or “SSS.” Raw time series sensor data, total suspended solids, and depth data from SSS were published at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566. A subset of data from the SSS samples were published in the contiguous United States (CONUS)-Scale Model-Sample (CM) study data package available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689 that presents data from across the CONUS. They include dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), total nitrogen (TN), grain size, aerobic sediment respiration, dissolved oxygen (DO), and temperature. Parent IDs and Site IDs are consistent between the SSS and CM data packages, and they can be mapped directly so data across packages can be used together. Field metadata for the samples in this da This dataset is comprised of one main data folder with four subfolders. The main data folder contains of (1) file-level metadata; (2) data dictionary; (3) total/water column/sediment respiration; (4) gross primary production (GPP); (5) median grain size (d50); and (6) annual oxygen consumption. The “Figures” subfolder contains the figures used in the paper and all intermediate files (including geospatial files). The “Published_Data” contains a readme directing the user to download the public data to reproduce analyses and figures. The “Scripts” folder contains all scripts used in the analyses that were not part of running StreamMetabolizer. Lastly, the “Stream_Metabolizer” folder contains all files associated with running StreamMetabolizer including (1) model input files, (2) model output files, (3) processing scripts, (4) histogram plots of the outputs, and (5) an R project. All files are .csv, .pdf, .R, .Rmd, .Rproj, .html, .png, .txt, .qgz, .cpg, .dbf, .prj, .shp, .shp.ea.iso.xml, .shp.iso.xml, .shx, .sbn. ta package can be found at either link. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

CaloFlow for CaloChallenge dataset 1

CALOFLOW is a new and promising approach to fast calorimeter simulation based on normalizing flows. Applying CALOFLOW to the photon and charged pion ≥ant showers of Dataset 1 of the Fast Calorimeter Simulation Challenge 2022, we show how it can produce high-fidelity samples with a sampling time that is several orders of magnitude faster than ≥ant. We demonstrate the fidelity of the samples using calorimeter shower images, histograms of high level features, and aggregate metrics such as a classifier trained to distinguish CALOFLOW from ≥ant samples.

Physics↗