Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Correlation-aware binning for small-angle neutron scattering via Gaussian-process inference

Binning in small-angle neutron scattering (SANS) is typically performed empirically, with fixed parameters chosen for convenience rather than statistical optimality. Such practices often fail to balance statistical precision and spatial resolution, leading to inconsistencies across instruments and datasets. Here we establish a correlation-aware framework that determines the optimal bin width from first principles by extending the classical Freedman–Diaconis (FD) rule to account for inter-bin correlations with a Gaussian process. In this formulation, the scattering intensity is treated as a smooth stochastic field whose statistical coherence is described by a covariance matrix. Analytical expressions of errors derived from this model yield closed-form criteria that separate the total deviation into contributions from counting noise, aliasing distortion and curvature-dependent correlation effects. Expressed in reduced variables, the resulting dimensionless error surface reveals a continuous transition from the uncorrelated FD regime to the correlation-dominated limit, providing a unified description of noise suppression and resolution control. Because the formulation depends only on the profile characteristics of scattering intensity I(Q), specifically its average intensity and first- and second-order derivatives, it applies generally to any SANS measurement regardless of sample, instrument or geometry. Experimental validation using small- and ultra-small-angle neutron scattering data confirms the predicted scaling behavior, demonstrating that correlation-aware inference systematically reduces mean-squared error and enables information-efficient reproducible data reduction across materials and instruments.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

Lorentz factor for time-of-flight neutron Bragg and total scattering

We report the three fundamental origins of the Lorentz factor for neutron time-of-flight powder diffraction are revisited. A detailed derivation of the Lorentz factor is presented in the context of diffuse scattering modelling in reciprocal space when perfect periodicity is assumed, and the total scattering pattern is constructed in its discrete form – the factor in this case becomes 1/Q 2 (or d 2 ). Discussion is also presented with respect to practical data reduction where a vanadium measurement is usually taken as the normalization factor (to account for various factors such as detector efficiency), and it is shown that the existence of the Lorentz factor is independent of such a normalization process.

36 MATERIALS SCIENCE↗

SZ3: A Modular Framework for Composing Prediction-Based Error-Bounded Lossy Compressors

Today's scientific simulations require a significant reduction of data volume because of extremely large amounts of data they produce and the limited I/O bandwidth and storage space. Here, error-bounded lossy compression has been considered one of the most effective solutions to the above problem. In practice, however, the best-fit compression method often needs to be customized or optimized in particular because of diverse characteristics in different datasets and various user requirements on the compression quality and performance. In this paper, we address this issue with a novel modular, composable compression framework named SZ3. Our contributions are four-folds. (1) We develop SZ3 which features an innovative modular abstraction for the prediction-based compression framework, such that compression modules can be plugged in easily to create new compressors based on characteristics of data and user requirements. (2) We create a new compression pipeline by SZ3 for GAMESS data, which significantly improves the compression ratios over state-of-the-art compressors. (3) We develop an adaptive compression pipeline by SZ3 for APS data with minimal efforts, which leads to the best rate-distortion among all existing error-bounded lossy compressors for any bit-rate. (4) We compare the sustainability of SZ3 with leading error-bounded prediction-based compressors, and then demonstrate the necessity of diverse pipelines by integrating and evaluating several compression pipelines on diverse scientific datasets from multiple disciplines. Experiments show that SZ3 incurs very limited overhead in compressor integration and our customized compression pipelines lead to up to 20% improvement in compression ratios under the same data distortion, when compared with the best existing approach.

97 MATHEMATICS AND COMPUTING↗

DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained Deduplication

Log-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSM-tree-based key-value store deployments. Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness. In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads. DedupKV features three key innovations: (1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity. Additionally, dynamic granularity management reduces deduplication metadata overhead. We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment. Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2 ×, 2 ×, and 1.6 × for real KV datasets.

Jamil, Safdar [Sogang University]↗

Enabling Command-and-Control in Advanced In Situ Workflows

Scientific discovery is progressing towards autonomous science with the combination of scientific instruments, high-performance computing, and artificial intelligence in complex workflows. This evolution introduces new requirements for managing scientific workflows, including feedback loops, near real-time constraints, and the ability to dynamically control workflow execution. In situ workflows that analyze and visualize data as it is generated are well-suited to satisfy stringent time constraints and their iterative nature offers greater opportunities for command-and-control. However, only a few of the many workflow management systems available have been specifically designed to manage in situ workflows and often lack support for automated feedback loops that allow analysis and visualization components to interact with the main scientific data producer. To address this need, we present in this paper how to add command-and-control capabilities to a workflow management system. We identify the functional design requirements of such a command-and-control system, detail its architecture, interface, and core mechanisms, and illustrate how advanced in situ workflows can leverage command-and-control in three use cases: graceful termination with checkpoint, dynamic and adaptive data reduction, and event-triggered analysis.

Mehta, Kshitij [ORNL] (ORCID:0000000297149981)↗

pySimpleMask

SF-26-118 pySimpleMask is a tool for creating masks and Q-partition maps for X-ray scattering patterns, supporting SAXS, WAXS, and XPCS data reduction. It ships both a desktop GUI and a headless Python API that can drive the full pipeline from scripts.

Chu, Miaoqi [Argonne National Laboratory (ANL), Ar↗

ZFP: A compressed array representation for numerical computations

HPC trends favor algorithms and implementations that reduce data motion relative to FLOPS. We investigate the use of lossy compressed data arrays in place of traditional IEEE floating point arrays to store the primary data of calculations. Simulation is fundamentally an exercise in controlled approximation, and error introduced by finite-precision arithmetic (or lossy compression) is just one of several sources of error that need to be managed to ensure sufficient accuracy in a computed result. We describe ZFP, a compressed numerical format designed for in-memory storage of multidimensional arrays, and summarize theoretical results that demonstrate that the error of repeated lossy compression can be bounded and controlled. Furthermore, we establish a relationship between grid resolution and compression-induced errors and show that, contrary to conventional floating point, ZFP reduces finite-difference errors with finer grids. We present example calculations that demonstrate data reduction by 4x or more with negligible impact on solution accuracy. Our results further demonstrate several orders-of-magnitude increase in accuracy using ZFP over IEEE floating point and Posits for the same storage budget.

Lindstrom, Peter↗

Uncertainty Propagation from Experiment Measurements to Modeling Approaches: A Case for SMR Steam Entrainment Testing

To license new and advanced reactor designs, regulators must be convinced that their unique safety cases—relative to existing large scale reactors—have been adequately addressed by the designed reactor protection systems. In water cooled small modular reactors (SMRs), droplet entrainment in steam flow has significant implications on the progression of accident scenarios due to its compact design features, which requires representative test data applicable to SMR designs. Computer code, modeling and simulation (M&S) tools and models require adequate verification, assessment, and qualification. This includes M&S results validation against scaled empirical data within allowable uncertainty bands to gain regulatory approvals during the various stages of reactor system design, demonstration, and commercialization. However, measurement uncertainty within the empirical datasets and test data applicability ranges requires careful consideration of M&S inputs (i.e., boundary conditions, and initial conditions), and verification and validation efforts. This study focuses on uncertainty quantification in designing scaled test facilities for SMR applications with appropriate measurements and a standard data-reduction method to estimate thermal hydraulics characteristics parameters that incorporate physics phenomena of interest. In addition, this study supports the evaluation model development and assessment process using M&S that interfaces with advanced computing tools and digital twin capabilities. This will allow synchronization between experiment and modeling approaches for droplet entrainment testing and analysis, improving diagnostics, prognostics, and decision-making to accelerate regulatory approval.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Transfer learning for metamaterial design and simulation

Abstract We demonstrate transfer learning as a tool to improve the efficacy of training deep learning models based on residual neural networks (ResNets). Specifically, we examine its use for study of multi-scale electrically large metasurface arrays under open boundary conditions in electromagnetic metamaterials. Our aim is to assess the efficiency of transfer learning across a range of problem domains that vary in their resemblance to the original base problem for which the ResNet model was initially trained. We use a quasi-analytical discrete dipole approximation (DDA) method to simulate electrically large metasurface arrays to obtain ground truth data for training and testing of our deep neural network. Our approach can save significant time for examining novel metasurface designs by harnessing the power of transfer learning, as it effectively mitigates the pervasive data bottleneck issue commonly encountered in deep learning. We demonstrate that for the best case when the transfer task is sufficiently similar to the target task, a new task can be effectively trained using only a few data points yet still achieve a test mean absolute relative error of 3 % with a pre-trained neural network, realizing data reduction by a factor of 1000.

Peng, Rixi↗

TEAMER: Water Tunnel Data from Testing the Pterofin Skimmer Concept

Pterofin's Skimmer concept relies on a flapping and pitching hydrofoil to extract hydrokinetic energy from water flows. The concept aims to utilize unsteady fluid dynamics phenomena (added mass, shed vorticity, and unsteady boundary layer development) to achieve higher lift coefficients, enabling increased power density of the hydrokinetic device and a fundamental shift in the rpm/torque scaling of the power take off compared with turbines. The Applied Research Laboratory at Penn State, in collaboration with Pterofin, designed and built a proof-of-concept flapping/pitching mechanism which was subsequently tested in ARL's 12-inch water tunnel facility. The mechanical power supplied to or extracted from the mechanism was measured for a range of hydrofoils provided by Pterofin over operating conditions including reduced frequency, Reynolds number, and the ratio between pitching and flapping amplitudes. The power lost to friction in the mechanism was removed from the net power measurement by means of a bare hub tare, with the resultant hydrodynamic power being used to calculate a mechanism-independent and non-dimensional power coefficient. The product of this effort is a dataset describing the power coefficient of a hydrofoil having simultaneous pitching and flapping motions, both of which are approximately sinusoidal. Power coefficients were collected for a range of primary design variables including: - Reduced frequency: 0.01 to 0.95 - Pitching/flapping peak angle ratio: 1.5 to 3.0 - Chord-based Reynolds number: 60,000 to 560,000 Secondary design variables relating to the hydrofoil geometry were explored including: - Aspect ratio - Planform shape - Section thickness distribution - Hydrofoil position relative to the pitching axis - Hydrofoil sweep angle relative to the pitching axis Measured data are provided in mean and time series formats. MATLAB scripts are provided which can be used to generate figures of time-averaged and phase-averaged hydrodynamic power coefficients calculated from the measured data. A complete description of the experiment and data reduction can be found in the Post Access Report for the Pterofin Skimmer test effort which will be available on the TEAMER website. This work was supported by the Pacific Energy Ocean Trust via a TEAMER award.

16 TIDAL AND WAVE POWER↗

2019 Quantum Materials Young Investigators Workshop Report

The fourth edition of the workshop “Quantum Materials Young Investigators” was held June 6-7, 2019 at the Oak Ridge National Laboratory (ORNL). This workshop followed up on previous meetings that took place at ORNL starting in 2016. This important series of annual workshops, organized by Adam Aczel and Stuart Calder, ensures continuous engagement with the early career principal investigators (PIs) of the quantum materials neutron scattering community. Through a series of short scientific talks presented by the external PIs, members of the Neutron Scattering Division (NSD) learn about their scientific interests and needs. Presentations from ORNL staff members also help to inform the external PIs about the current state of ORNL’s quantum materials neutron scattering program. Topics covered include the current instrument suites, sample environment, data reduction, visualization, and analysis, and future neutron scattering and complementary capabilities. Feedback is solicited from the external PIs in all these key areas and communicated to NSD management in the form of a workshop report. This workshop series also helps to foster collaborations between NSD scientists and the research groups of the external PIs and plays an important role in attracting new neutron scattering users.

43 PARTICLE ACCELERATORS↗

2022 Quantum Materials Young Investigators Workshop Report

The fifth edition of the workshop “Quantum Materials Young Investigators” was held August 18-19, 2022 at the Oak Ridge National Laboratory (ORNL). This workshop followed up on previous meetings that took place at ORNL starting in 2016. This important series of annual workshops, organized by Adam Aczel and Stuart Calder, ensures continuous engagement with the early career principal investigators (PIs) of the quantum materials neutron scattering community. Through a series of short scientific talks presented by the external PIs, members of the Neutron Scattering Division (NSD) learn about their scientific interests and needs. Presentations from ORNL staff members also help to inform the external PIs about the current state of ORNL’s quantum materials neutron scattering program. Topics covered include neutron scattering instruments, sample environment, neutron polarization, data reduction, visualization, and analysis, and future neutron scattering and complementary capabilities. Feedback is solicited from the external PIs in some or all of these key areas and communicated to NSD management in the form of a workshop report. This workshop series also helps to foster collaborations between NSD scientists and the research groups of the external PIs and plays an important role in attracting new neutron scattering users.

36 MATERIALS SCIENCE↗

Extracting LANSCE Macrobunch Charge from WNR Fast Pickoff [Slides]

Over several days around December 18, 2023, high-charge minipulses and macropulses were sent to Target 2, also known as the Blue Room. This report presents the calibration of a stripline-type current monitor known as the Fast Pickoff, and subsequent charge data for many of these shots. This effort utilized a fast oscilloscope, a nearby Bergoz current monitor, and some data-reduction techniques.

43 PARTICLE ACCELERATORS↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Radiation-Hard Smart-Pixel Detector ASIC ReadOut with Digital AI in 28nm

Detectors at future high energy colliders will face enormous technical challenges. Disentangling the unprecedented numbers of particles expected in each event will require highly granular silicon pixel detectors with billions of readout channels. With event rates as high as 40 MHz, these detectors will generate petabytes of data per second. To enable discovery within strict bandwidth and latency constraints, future trackers must be capable of fast, power efficient, and radiation hard data-reduction at the source. This effort is pursuing the co-design development of high-performance readout smart pixel ASICs for a future Phase III High Luminosity upgrade of the Large Hadron Collider. A 1.6mm2 ASIC prototype was designed by Fermilab in CMOS 28 nm bulk process and submitted for manufacturing in February 2024. It leverages the analog front-end pixel design of a previous prototype fabricated and tested in 2023, which achieved a simulated detection level of ~400e- with 30fF input capacitance. The ROIC consists of two matrices of 16×16 smart pixels, each 25×25 μm2 in size. Each smart pixel contains a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. There is digital space for the integration of our fully combinatorial AI that performs momentum classification at the bunch crossing rate. The total power consumption is ∼6μW per pixel, which corresponds to ~1mW/cm2. The ASIC incorporates programmable front-end charge injection circuitry to generate pixel cluster charges during characterization. The cluster profile will be generated to duplicate hit characteristics of various momentum (pT). We will present early results from chip testing.

Parpillon, Benjamin↗

10 Boron-film-based gas detectors at ESS

The development of detectors for the European Spallation Source is an important parallel element to the effort put into the design and construction of the neutron source and instruments. The 10Boron-film-based detector developments that started over a decade ago as basic detector concepts have now reached technical maturity after intense prototyping work and numerous testing campaigns. Several of the ESS beamlines that will soon enter the commissioning phase started welcoming the detector systems built with the 10 B-film converter technology. The real-size demonstrators evolved in the last 2–3 years into a diverse suite of gas proportional counters for use in diffraction, reflectometry and small-angle scattering studies in spite of the operational complexity posed by the requirements for large-area coverage, low material budget and robustness for operation in the high-flux and high-radiation environment of ESS. The detectors are now being shipped to ESS or undergoing the last performance and calibration tests before being handed over to the instrument teams for installation in the host beamlines. A common feature of these detector systems is the very large number of readout channels that they are instrumented with in order to fulfill the demanding requirements for sensitivity, spatial resolution and count-rate capability. Across all of the detector types that will operate at ESS, the integration, testing and commissioning of the read-out technologies and software tools for data reduction, calibration and analysis are the focus of the detector and integration teams. These will yield a wealth of knowledge about their operation as well as initial results on the in-situ performance, a very important asset to ensure rapid commissioning of the detectors when neutrons from the ESS source become available. In this paper we will give an overview of the 10 B-film-based detector technologies included in the ESS detector suite that are currently facing the transition from the production phase to installation and integration with the other beamline components, which comes with its own specific challenges of both organizational and technical nature.

47 OTHER INSTRUMENTATION↗

The Third Data Release of the KODIAQ Survey

We present and make publicly available the third data release (DR3) of the Keck Observatory Database of Ionized Absorption toward Quasars (KODIAQ) survey. KODIAQ DR3 consists of a fully reduced sample of 727 quasars at 0.1 < z {sub em} < 6.4 observed with the Echellette Sepctrograph and Imager at moderate resolution (4000 ≤ R ≤ 10,000). DR3 contains 872 spectra available in flux calibrated form, representing a sum total exposure time of ∼2.8 megaseconds. These coadded spectra arise from a total of 2753 individual exposures of quasars taken from the Keck Observatory Archive (KOA) in raw form and uniformly processed using a data reduction package made available through the XIDL distribution. DR3 is publicly available to the community, housed as a higher level science product at the KOA and in the igmspec database.

74 ATOMIC AND MOLECULAR PHYSICS↗