Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

User Guide for Sample Reduction at GP-SANS

This manual is intended as a quick guide for data reduction of GP-SANS data. It includes all necessary steps to do the data reduction based on absolute calibration using the open beam method and how to transfer the reduced data to the personal computer system. If any errors are coming up so that the reduction script is not functioning as intended, please contact the instrument scientist.

97 MATHEMATICS AND COMPUTING↗

Another Set of Python Tools for Visualizing and Manipulating Small-Angle Neutron Scattering Data: Descriptions and Examples

The GP-SANS, Bio-SANS and EQ-SANS instruments at ORNL utilize drtsans for data reduction. drtsans is built on Python, and it can be run using python scripts and Jupyter notebooks. The flexibility afforded by Python makes it possible to incorporate additional actions into the scripts used for data reduction, such as analysis and visualization. Here, a new set of tools for visualizing and manipulating SANS data that can be incorporated into the data reduction scripts for the ORNL SANS instruments, or employed during post–processing, is presented that expands the capabilities of the two previously-released tool sets.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sensor Co-design for $\textit{smartpixels}$

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in the first level of the trigger for a hadron collider. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p$_T$) based on the geometrical shape of the charge deposition (``cluster''). To design a viable detector for deployment at an experiment, the dependence of the NN as a function of the sensor geometry, external magnetic field, and irradiation must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. A smaller sensor pitch in the bending direction improves the p$_T$ discrimination, but a larger pitch can be partially compensated with detector depth. An external magnetic field parallel to the sensor plane induces Lorentz drift of the electron-hole pairs produced by the charged particle, broadening the cluster and improving the network performance. The absence of the external field diminishes the background rejection compared to the baseline by $\mathcal{O}$(10%). Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by $\sim$ 30 - 60%, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data.

Shekar, Danush [Illinois U., Chicago]↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

200-IA-1 Operable Unit Human Health Risk and Kd Screen

The purpose of this environmental calculation file (ECF) is to provide the following: Document the data processing and data reduction steps taken to prepare the 200-IA-1 Operable Unit (OU) data set that will be used to calculate the sample-specific screening level human health risk evaluation; Document the data processing and data reduction steps taken to prepare the 200-IA-1 OU data set that will be used to identify analytes that could potentially impact groundwater in the future beneath the 200-IA-1 OU representative waste sites; Document the assumptions, equations, and methodologies used to calculate the screening levels for human health cancer risks and noncancer hazards for each representative waste site assigned to the 200-IA-1 OU. Individual measured soil concentrations from 0 to 4.6 m (15 ft) below ground surface (bgs) (shallow vadose zone) are used to calculate the total excess lifetime cancer risk (ELCR) and hazard index (HI) for the outdoor worker scenario to determine if there is a basis for remedial action. Individual measured soil concentrations from the ground surface to the groundwater table are used for the distribution coefficient (Kd) screen to identify analytes that could potentially impact groundwater in the future beneath the 200-IA-1 OU representative waste sites. This ECF supports DOE/RL-2020-51, 200-IA-1 OU Focused Feasibility Study, under the Comprehensive Environmental Response, Compensation, and Liability Act of 1980 (CERCLA). A risk characterization based upon the evaluation of the health risk estimates developed in this ECF will be presented in the focused feasibility study report.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Adaptive Methods for Radial Basis Functions

Radial basis functions (RBFs) are a powerful tool for constructing high-order accurate reduced representations of scattered data in arbitrary dimension and on manifolds. We present a method of constructing data approximations in which we utilize a functional tail to capture a global background profile and a RBF neural network (NN) to capture the smaller-scale features. In the RBF NN the RBF centers, matrix shape parameters were selected adaptively for each RBF. We also utilized a geodesic notion of distance on the manifold on which the data lies, e.g., the spherical geodesic for data on the sphere. Although each of these ideas have been been investigated separately in previous works, their combination into a single algorithm is novel. We defined a machine learning problem in which these properties are learned to minimize the data reduction error. We demonstrate the algorithm for applications of scattered data reduction in the plane and on the sphere.

97 MATHEMATICS AND COMPUTING↗

Classical and quantum compression for edge computing: the ubiquitous data dimensionality reduction

Edge computing aims to address the challenges associated with communicating and transferring large amounts of data generated remotely to a data center in a timely and efficient manner. A central pillar of edge computing is local (i.e., at- or near-source) data processing capability so that data transfer to a data center for processing can be minimized. Data compression at the edge is therefore a natural component of edge workflows. Here we present a survey of data compression algorithms with a focus on edge computing. Not all compression algorithms can accommodate the data type heterogeneity, tight processing and communication time constraints, or energy efficiency requirement characteristics of edge computing. We discuss specific examples of compression algorithms that are being explored in the context of edge computing. We end our review with a brief survey of emerging quantum compression techniques that are of importance in quantum information processing, including the proposed concept of quantum edge computing.

97 MATHEMATICS AND COMPUTING↗

Data Release 1 of the Dark Energy Spectroscopic Instrument

In 2021 May the Dark Energy Spectroscopic Instrument (DESI) collaboration began a 5 yr spectroscopic redshift survey to produce a detailed map of the evolving three-dimensional structure of the Universe between z = 0 and z ≈ 4. DESI’s principal scientific objectives are to place precise constraints on the equation of state of dark energy, the gravitationally driven growth of large-scale structure, and the sum of the neutrino masses, and to explore the observational signatures of primordial inflation. We present DESI DR1, which consists of all data acquired during the first 13 months of the DESI main survey, as well as a uniform reprocessing of the DESI Survey Validation data, which were previously made public in the DESI Early Data Release. The DR1 main survey includes high-confidence redshifts for 18.7M objects, of which 13.1M are spectroscopically classified as galaxies, 1.6M as quasars, and 4M as stars, making DR1 the largest sample of extragalactic redshifts ever assembled. We summarize the DR1 observations, the spectroscopic data-reduction pipeline and data products, large-scale structure catalogs, value-added catalogs, and describe how to access and interact with the data. In addition to fulfilling its core cosmological objectives with unprecedented precision, we expect DR1 to enable a wide range of transformational astrophysical studies and discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

Integrating ORNL’s HPC and Neutron Facilities with a Performance-Portable CPU/GPU Ecosystem

We explore the development of a performance-portable CPU/GPU ecosystem to integrate two of the US Department of Energy’s (DOE’s) largest scientific instruments, the Oak Ridge Leadership Computing facility and the Spallation Neutron Source (SNS), both of which are housed at Oak Ridge National Laboratory. We select a relevant data reduction workflow use-case to obtain the differential scattering cross-section from data collected by SNS’s CORELLI and TOPAZ instruments. We compare the current CPU-only production implementation using the Garnet Python multiprocess package based on the Mantid C++ framework against our proposed CPU/GPU implementation that uses the LLVM-based, just-in-time Julia scientific language and the JACC.jl performance-portable package. Two proxy apps were developed: (i) an app for extracting relevant Mantid kernels (MDNorm) in C++ and (ii) the Julia MiniVATES.jl miniapp. We present performance results for NVIDIA A100 and AMD MI100 GPUs and AMD EPYC 7513 and 7662 CPUs. The results provide insights for future generations of data reduction software that can embrace performance portability for an integrated research infrastructure across DOE’s experimental and computational facilities.

Hahn, Steven↗

More Tools for Visualization and Analysis of Small-Angle Neutron Scattering Data: Descriptions and Examples

With the adoption of drtsans as the data reduction software for the GP-SANS, Bio-SANS and EQ-SANS instruments at ORNL, tools for data visualization and analysis that can be integrated into drtsans scripts are needed to further improve the user experience. New tools that do not need to be incorporated directly into data reduction scripts can also positively impact users during their experiments. In this report, a new set of tools is presented that complements the previous set released. The set includes tools for both fitting data and for visualizing data.

42 ENGINEERING↗

Uniform-in-phase-space data selection with iterative normalizing flows

Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that are routinely generated. In applications that are constrained by memory and computational intensity, excessively large datasets may hinder scientific discovery, making data reduction a critical component of data-driven methods. Datasets are growing in two directions: the number of data points and their dimensionality. Whereas dimension reduction typically aims at describing each data sample on lower-dimensional space, the focus here is on reducing the number of data points. A strategy is proposed to select data points such that they uniformly span the phase-space of the data. The algorithm proposed relies on estimating the probability map of the data and using it to construct an acceptance probability. An iterative method is used to accurately estimate the probability of the rare data points when only a small subset of the dataset is used to construct the probability map. Instead of binning the phase-space to estimate the probability map, its functional form is approximated with a normalizing flow. Therefore, the method naturally extends to high-dimensional datasets. The proposed framework is demonstrated as a viable pathway to enable data-efficient machine learning when abundant data are available.

97 MATHEMATICS AND COMPUTING↗

SNAPRed: Reduction of multidimensional neutron time-of-flight diffraction data

SNAP is a neutron time-of-flight diffractometer at the Spallation Neutron Source operated by Oak Ridge National Laboratory. It generates large arrays of neutron detection events that encode the crystalline atomic structure of materials under study. SNAPRed is an application that makes these datasets accessible to end users by orchestrating the process of data reduction while automatically managing the variable neutron instrumentation configuration. It supports arbitrary grouping and masking of individual detector pixels and includes custom-developed data compression approaches to accommodate the large volumes of data generated by the SNAP instrument.

Diffraction↗

Laser-Induced Spectrochemical Assay for Uranium Enrichment (LISA-UE)

Uranium hexafluoride (UF6) is the uranium compound typically involved in uranium enrichment process. As the first line of defense against nuclear proliferation, accurate determinations of the uranium enrichment ratio in UF6 are critical for materials verification, accounting and safeguards. Shipping gaseous UF6 samples off-site for analysis with mass spectrometry is cumbersome and costly, and results are not available for some time (months). In-field UF6 enrichment assay has the potential to substantially reduce the time, logistics and expense of sample handling. At present, COMPUCEA is the only accepted method for UF6 enrichment assay in the field. Laser-Induced Spectrochemical Assay for Uranium Enrichment (LISA-UE) is an all-optical (based on laser induced plasma emission) analytical technique intended for fieldable, accurate, precise and rapid UF6 enrichment assay. In its operation, laser induced plasma is created directly in the gaseous UF6 sample. Because different U isotopes emit at slightly different wavelengths, the isotopic information of the UF6 sample is inherently encoded in the atomic emission from the plasma. Isotopic emissions from 235U and 238U are measured simultaneously, which eliminate correlated noise from the laser induced plasma. Isotopic information of the UF6 sample can be extracted from the acquired spectrum with theoretical multi-variable non-linear spectral fitting. To date, advances made by the LISA-UE research team include optimization of the spectral window for direct gaseous UF6 enrichment assay with laser induced plasma, development of data reduction algorithms, and demonstrations of the LISA-UE technique with gaseous UF6 samples. In this presentation, the technical aspect of LISA-UE will be overviewed, the data reduction algorithm will be described, and performance of the technique will be discussed.

Chan, George↗

Developing Methodology to Determine Pu Isotopic Composition by Laser Ablation MC-ICP-MS

This project will develop methodology to analyze the Pu isotope ratio in mixed U-Pu particles by laser ablation MC-ICP-MS. This will involve: 1) testing and validation of the Pu analytical method using mixed U-Pu solutions and Pu doped glasses, 2) isotopic analysis of mixed U-Pu particles by laser ablation MC-ICP-MS, and 3) continued development of the R-based data reduction program. Details regarding the analytical method development and results of QC testing will be output as a deliverable to the IAEA, along with an updated version of the LARA data reduction software.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

SonicPy: a suite of programs for ultrasound pulse-echo data acquisition and analysis

Sound speed and elastic constants measurements in solids and liquids are commonly performed using the ultrasound pulse-echo technique. Recent advances have expanded the use of this technique at numerous high pressure synchrotron beamlines and offline laboratories. However, the increased experimental throughput has revealed many limitations in existing software for handling the rapid measurement and the subsequent data-reduction. Here, we report the development of a collection of computer programs for sound speed measurements using the ultrasound pulse-echo technique, compatible with stepped multi-frequency, as well as broadband-pulse, couplant-corrected methods. The programs provide a highly interactive graphical interface, enable efficient measurement, exploration and near real-time analysis of the ultrasound data, and contain features useful for working with samples under high pressure and/or high temperature. The included analysis programs can alleviate the time required for data reduction from hours to less than a minute, allowing users to make timely and informed decisions regarding the appropriate experimental parameters.

97 MATHEMATICS AND COMPUTING↗