Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA REDUCTION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator↗

Towards Autonomous Experiments by Connecting High Performance Microscopy with High Performance Computing

The digitization of controls, data, and analysis in microscopy is bringing the idea of autonomous microscopes closer to reality than ever before. Automated transmission electron microscopy (TEM) is already fairly routine for some experiments the only require simple repetitive tasks such as imaging biological macromolecules for single particle cryoEM [1], tilt series for electron tomography [2], and movies for crystallography [3]. The vast majority of TEM experiments are conducted completely by human operators who choose the regions of interest, optimize experimental parameters, and make decisions about data quality visually during an experiment. The field is still a long way from having completely autonomous TEMs that can adapt to sample difficulties and tune experimental parameters based on data quality and desired experimental outcomes. Part of the issue is the lack of capability for feeding information learned from on-line, live data analysis back into the on-going experiment [4]. Furthermore, this presentation will discuss current capabilities for large scale data reduction and analysis using high performance computing (i.e. supercomputing) and progress towards developing a true feed-back loop that places data analysis and theory in the experimental loop.

97 MATHEMATICS AND COMPUTING↗

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.↗

The S-PLUS Fifth Data Release: Over 4500 Square Degrees of the Southern Sky and a Multicolor View of the Hydra and Antlia Galaxy Clusters

Abstract We present the fifth data release (DR5) of the Southern Photometric Local Universe Survey (S-PLUS), covering 4592 deg 2 across 2491 fields. Observations were conducted with the T80-South, a Brazilian robotic telescope equipped with the Javalambre 12-filter system, containing five broadband and seven narrowband filters. Data products feature FITS images and extensive catalogs containing fluxes, magnitudes, and shape parameters for over 113 million detections. In addition, several value-added catalogs are provided, offering photometric redshifts (photo- zs ), object classifications, masks, and extinction coefficients. For the first time, this release includes coverage of 110 deg 2 along the Galactic disk, facilitating new research into Galactic structure and stellar populations. The release also provides full coverage of the Hydra Supercluster and numerous other nearby clusters with improved data reduction and calibration, enhancing photo- z accuracy, which is vital for large-scale structure studies. A preliminary analysis of the Hydra and Antlia galaxy clusters up to 5 × R 200 yields an updated catalog of 1706 cluster members based on both spectroscopic and high-quality photo- zs . Our photometric data shows that both clusters have a similar proportion of galaxies with an H α excess relative to their clustercentric distance, though Hydra has a higher fraction near its center. Additionally, the spatial distribution of all objects in our sample highlights a bridge connecting both clusters. We verify that S-PLUS DR5 provides a solid foundation for future scientific investigations, ranging from solar system studies to cosmology.

Vinicius Rodrigues de Lima, Erik [Universidade de ↗

Development of Parameters for the Particle Size Distribution of TATB

Laser light scattering (LLS), manual counting of scanning electron microscopy (SEM) images, and reduction of SEM images using ImageJ (open source program) to determine Feret (caliper) diameters were applied to determine particle size distribution (PSD) of four preparations of 1,3,5‐triamino‐2,4,6‐trinitrobenzene (TATB) and yttria‐stabilized zirconia (YSZ), an SEM certified standard. Mie theory was used to reduce the LLS data. The spherical nature of the YSZ made it a good candidate for LLS. Variations in n , the refractive index, and iκ, the imaginary component, produced very little change in the PSD. However, changing the carrier liquid from H 2 O to a 40% aqueous sucrose solution, thereby changing the carrier refractive index, n 0 , substantially affected the PSD. The Mie complex refractive indices for the YSZ were n = 2.200, iκ = 0.100, with a 40% aqueous sucrose solution, n 0 = 1.400. The triclinic crystal structure of TATB made refractive index determinations more difficult, so a study was conducted varying Mie parameters and comparing them to the same data reduced using the Fraunhofer theory. Changing the n and κ parameters produced PSD with a small concentration of particles less than 1 µm in size or none in this range. SEM images, Feret data, manual counting, and Fraunhofer data reduction indicate particles less than 1 µm are probably < 5% in concentration. The final selection of Mie parameters for TATB was n = 2.283, iκ = 0.1, and suspension medium, n 0 = 1.330. Finally, computations, using density functional theory produced similar parameters.

Feret diameter↗

In-pixel integration of signal processing and AI/ML based data filtering for particle tracking detectors

We present the first physical realization of in-pixel signal processing with integrated AI-based data filtering for particle tracking detectors. Building on prior work that demonstrated a physics-motivated edge-AI algorithm suitable for ASIC implementation, this work marks a significant milestone toward intelligent silicon trackers. Our prototype readout chip performs real-time data reduction at the sensor level while meeting stringent requirements on power, area, and latency. The chip is taped-out in 28nm TSMC CMOS bulk process, which has been shown to have sufficient radiation hardness for particle experiments. This development represents a key step toward enabling fully on-detector edge AI, with broad implications for data throughput and discovery potential in high-rate, high-radiation environments such as the High-Luminosity LHC.

Parpillon, Benjamin [Fermilab; Illinois U., Chicag↗

tobac v1.5: introducing fast 3D tracking, splits and mergers, and other enhancements for identifying and analysing meteorological phenomena

There is a continuously increasing need for reliable feature detection and tracking tools based on objective analysis principles for use with meteorological data. Many tools have been developed over the previous 2 decades that attempt to address this need but most have limitations on the type of data they can be used with, feature computational and/or memory expenses that make them unwieldy with larger datasets, or require some form of data reduction prior to use that limits the tool's utility. The Tracking and Object-Based Analysis of Clouds (tobac) Python package is a modular, open-source tool that improves on the overall generality and utility of past tools. A number of scientific improvements (three spatial dimensions, splits and mergers of features, an internal spectral filtering tool) and procedural enhancements (increased computational efficiency, internal regridding of data, and treatments for periodic boundary conditions) have been included in tobac as a part of the tobac v1.5 update. These improvements have made tobac one of the most robust, powerful, and flexible identification and tracking tools in our field to date and expand its potential use in other fields. Future plans for tobac v2 are also discussed.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML) and silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

Gonski, Julia [SLAC]↗

A 28 nm multiply-accumulate ASIC architecture for on-chip data compression in MHz frame rate X-ray and electron pixel detectors

Modern X-ray detector systems urgently require compact, efficient, and fast data compression schemes to handle the transmission of big data from pixel arrays, enabling frame rates in the MHz regime. Here, in this work, a data compression ASIC that implements a streaming fixed-length lossy compression scheme is introduced and analyzed, proving the feasibility and benefits of on-chip compression. The compression scheme utilizes a vector matrix product logic, which performs a number of floating-point multiplications, additions, and accumulations. The logic is verified, synthesized, and shown to fit in the area resource available for the X-ray detector under study, which comprises 192 × 168 pixels each of 12-bit width, and having a total area of 20 mm× 20 mm, about 2 mm× 20 mm of which are available for the digital logic. Several system architectures, precisions, and compression ratios ranging from 100 to 250 were analyzed to pave the way for on-chip fixed-length compression (e.g., principal component analysis, singular value decomposition) and data reduction (e.g., azimuthal integration) for X-ray and electron detectors.

Data compression↗

ClassNMSW- a real-time classification approach for non-recycled municipal solid waste using hyperspectral imaging

Real-time classification of non-recycled municipal solid waste (NMSW) is essential for efficient valorization. This study introduces ClassNMSW, a comprehensive framework for classifying 22 NMSW subclasses under industrial constraints by using hyperspectral imaging (HSI). A primary innovation of this work is the development of a variance-controlled spectral extraction algorithm. Unlike traditional methods that rely on simple averaging, this approach systematically investigates the extent of pixel extraction to minimize the loss of critical chemical information while maximizing data reduction thus ensuring high spectral fidelity with low computational cost. The approach developed in this work integrates automated, computer-vision-based background removal, eliminating the need for the manual thresholding common in current literature. To resolve ambiguities among chemically similar subclasses, a tiered classification and multi-camera fusion strategy (NIR17 and NIR22) is implemented. Results demonstrate that ClassNMSW achieves an object-wise weighted accuracy of 98.70% for single-sensor configurations and 100% under sensor fusion. A novel rolling-window strategy satisfies desired end-to-end latency of <2 s, satisfying the strict deterministic requirements of high-speed industrial sorting environments. The ClassNMSW framework provides a scalable foundation for advancing circularity and resource recovery in large-scale waste valorization operations.

99 - GENERAL AND MISCELLANEOUS↗

Investigating Quantum Materials with Half-Polarized Diffraction and magnetic PDF analysis at the HB-2A Neutron Powder Diffractometer

Local magnetic ordering and anisotropy is often central to the emergent behavior and subsequent functional properties in quantum materials and beyond. Neutron powder diffraction provides a straightforward yet extremely powerful technique for quantitative measurements of microscopic magnetic properties. The HB-2A powder diffractometer located at the High Flux Isotope Reactor in ORNL is traditionally utilized for long-range magnetic structure determination. Recently these capabilities have been extended to include methods aimed at accessing local magnetism: Half- polarized neutron powder diffraction (pNPD) and magnetic pair distribution function (mPDF) analysis. These two distinct techniques are possible on HB-2A due to the versatility of the instrument’s reciprocal space coverage, resolution and novel ultra-low temperature multi-sample changers that operate down to dilution refrigerator temperatures. This provides unique capabilities not found on any powder diffraction instrument and is particularly well suited to investigations of magnetic quantum materials. The development and implementation of these techniques will be discussed with a series of science case examples ranging from geometric frustrated magnets to magnetic metal-organic frameworks. Data reduction and analysis tools will be presented that enable the extraction of the local site susceptibility tensor and local spin-spin correlations in real space. Finally, potential combinations of these techniques in the form of half-polarized magnetic pair distribution function (pmPDF) analysis will be considered. Looking forward, HB-2A is undergoing a detector upgrade that will be in the user program by 2026. This will offer an order of magnitude increase in count rates to further aid the development of these often low signal measurements and provide new scientific capabilities.

Neutron Scattering↗

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Simulations of Heat Transfer Using Tight-Fitting Twisted Tape Inserts for First Wall Cooling in Molten Salt Breeder Blankets

One of the major components in fusion energy systems is the fusion blanket, which has a vacuum vessel to contain the plasma. As part of the fusion blanket/vacuum vessel, the first wall and plasma-facing components require sufficient cooling to prevent material degradation during operation from the superheated plasma. Most fusion blanket concepts involve first wall and divertor coolant channels with heat transfer enhancements (HTEs) that are intended to withstand the incident high heat fluxes of 1 to 5 MW/m 2 . Twisted tape inserts are a proposed HTE that have been investigated previously for first wall cooling and monoblock divertor cooling channels and in other nonfusion heat transfer components. By inserting twisted tapes into straight pipes, the amount of turbulence in the system can be increased at lower Reynolds numbers by swirling the flow. This results in better heat transfer characteristics with marginal increases in frictional pressure losses. In particular, simulations of high-Prandtl-number fluids such as the proposed molten salt FLiBe in twisted tapes, which is prototypic to liquid immersion blankets, have not been previously explored. Here, in this study, we simulate various Prandtl numbers in pipes with twisted tape inserts using large eddy simulations to determine the effects of increasing Prandtl numbers on heat transfer performance. The quantities of particular interest are the Nusselt number and the friction factor, which were recovered using data reduction techniques to determine impacts on heat transfer and pressure losses. This work serves as a starting point for determining the feasibility of twisted tape inserts for liquid immersion blanket concepts.

LES↗

The Simons Observatory: science goals and forecasts for the enhanced Large Aperture Telescope

We describe updated scientific goals for the wide-field, millimeter-wave survey that will be produced by the Simons Observatory (SO). Significant upgrades to the 6-meter SO Large Aperture Telescope (LAT) are expected to be complete by 2028, and will include a doubled mapping speed with 30,000 new detectors and an automated data reduction pipeline. In addition, a new photovoltaic array will supply most of the observatory's power. The LAT survey will cover about 60% of the sky at a regular observing cadence, with five times the angular resolution and ten times the map depth of the Planck satellite. The science goals are to: (1) determine the physical conditions in the early universe and constrain the existence of new light particles; (2) measure the integrated distribution of mass, electron pressure, and electron momentum in the late-time universe, and, in combination with optical surveys, determine the neutrino mass and the effects of dark energy via tomographic measurements of the growth of structure at redshifts z ≲ 3; (3) measure the distribution of electron density and pressure around galaxy groups and clusters, and calibrate the effects of energy input from galaxy formation on the surrounding environment; (4) produce a sample of more than 30,000 galaxy clusters, and more than 100,000 extragalactic millimeter sources, including regularly sampled AGN light-curves, to study these sources and their emission physics; (5) measure the polarized emission from magnetically aligned dust grains in our Galaxy, to study the properties of dust and the role of magnetic fields in star formation; (6) constrain asteroid regoliths, search for Trans-Neptunian Objects, and either detect or eliminate large portions of the phase space in the search for Planet 9; and (7) provide a powerful new window into the transient universe on time scales of minutes to years, concurrent with observations from the Vera C. Rubin Observatory of overlapping sky.

79 ASTRONOMY AND ASTROPHYSICS↗

Smart pixel sensors: towards on-sensor filtering of pixel clusters with deep learning

Highly granular pixel detectors allow for increasingly precise measurements of charged particle tracks. Next-generation detectors require that pixel sizes will be further reduced, leading to unprecedented data rates exceeding those foreseen at the High- Luminosity Large Hadron Collider. Signal processing that handles data incoming at a rate of $\mathcal{O}$(40 MHz) and intelligently reduces the data within the pixelated region of the detector at rate will enhance physics performance at high luminosity and enable physics analyses that are not currently possible. Using the shape of charge clusters deposited in an array of small pixels, the physical properties of the traversing particle can be extracted with locally customized neural networks. In this first demonstration, we present a neural network that can be embedded into the on-sensor readout and filter out hits from low momentum tracks, reducing the detector's data volume by 57.1%–75.7%. The network is designed and simulated as a custom readout integrated circuit with 28 nm CMOS technology and is expected to operate at less than 300 μW with an area of less than 0.2 mm 2 . The temporal development of charge clusters is investigated to demonstrate possible future performance gains, and there is also a discussion of future algorithmic and technological improvements that could enhance efficiency, data reduction, and power per area.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Correlation-aware binning for small-angle neutron scattering via Gaussian-process inference

Binning in small-angle neutron scattering (SANS) is typically performed empirically, with fixed parameters chosen for convenience rather than statistical optimality. Such practices often fail to balance statistical precision and spatial resolution, leading to inconsistencies across instruments and datasets. Here we establish a correlation-aware framework that determines the optimal bin width from first principles by extending the classical Freedman–Diaconis (FD) rule to account for inter-bin correlations with a Gaussian process. In this formulation, the scattering intensity is treated as a smooth stochastic field whose statistical coherence is described by a covariance matrix. Analytical expressions of errors derived from this model yield closed-form criteria that separate the total deviation into contributions from counting noise, aliasing distortion and curvature-dependent correlation effects. Expressed in reduced variables, the resulting dimensionless error surface reveals a continuous transition from the uncorrelated FD regime to the correlation-dominated limit, providing a unified description of noise suppression and resolution control. Because the formulation depends only on the profile characteristics of scattering intensity I(Q), specifically its average intensity and first- and second-order derivatives, it applies generally to any SANS measurement regardless of sample, instrument or geometry. Experimental validation using small- and ultra-small-angle neutron scattering data confirms the predicted scaling behavior, demonstrating that correlation-aware inference systematically reduces mean-squared error and enables information-efficient reproducible data reduction across materials and instruments.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained Deduplication

Log-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSM-tree-based key-value store deployments. Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness. In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads. DedupKV features three key innovations: (1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity. Additionally, dynamic granularity management reduces deduplication metadata overhead. We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment. Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2 ×, 2 ×, and 1.6 × for real KV datasets.

Jamil, Safdar [Sogang University]↗

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)↗