Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Evaluation of Apex Alpha with LabWare-LIMS Data Reduction for Alpha Analyses

Savannah River Site's Environmental Bioassay Laboratory is migrating its laboratory information management system (LIMS) from SQL-LIMS to Oracle Labware-LIMS (LW-LIMS) systems to align with the Department of Energy's cyber security policies. Concurrently with the LIMS upgrade, the current VMS based Alpha Measurement System (AMS) software is being replaced with Apex Alpha software to reduce other cyber vulnerabilities, streamline procedural production aspects of data handling, and improve outdated reduction processes to align the program with requirements outlined in ANSI N13.30. The primary technical changes which are being implemented are: - Reagent contributions are added in data reduction, - Moving average background equivalent activity are used for gross signal corrections, - Measured uncertainty of both are included in the decision level and propagated uncertainty, - Tracer levels are increased for precision and accuracy improvements, and - Reports are modified for compliance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool for compressing large data matrices that arise in simulations of partial differential equations (PDEs), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. Here, to address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices in situ. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

Column subset selection↗

Efficient data reduction for time-of-flight neutron scattering experiments on single crystals

Event-mode data collection presents remarkable new opportunities for time-of-flight neutron scattering studies of collective excitations, diffuse scattering from short-range atomic and magnetic structures, and neutron crystallography. In these experiments, large volumes of the reciprocal space are surveyed, often using different wavelengths and counting times. These data then have to be added together, with accurate propagation of the counting errors. This paper presents a statistically correct way of adding and histogramming the data for single-crystal time-of-flight neutron scattering measurements. In order to gain a broader community acceptance, particular attention is given to improving the efficiency of calculations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybrid Approaches for Data Reduction of Spatiotemporal Scientific Applications

Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions.

Li, Xiao↗

Machine Learning Techniques for Data Reduction of Climate Applications

Scientists conduct large-scale simulations to compute derived quantities-of-interest (QoI) from primary data. Often, QoI are linked to specific features, regions, or time intervals, such that data can be adaptively reduced without compromising the integrity of QoI. For many spatiotemporal applications, these QoI are binary in nature and represent presence or absence of a physical phenomenon. We present a pipelined compression approach that first uses neural-network-based techniques to derive regions where QoI are highly likely to be present. Then, we employ a Guaranteed Autoencoder (GAE) to compress data with differential error bounds. GAE uses QoI information to apply low-error compression to only these regions. This results in overall high compression ratios while still achieving downstream goals of simulation or data collections. Experimental results are presented for climate data generated from the E3SM Simulation model for downstream quantities such as tropical cyclone and atmospheric river detection and tracking. These results show that our approach is superior to comparable methods in the literature.

Li, Xiao [University of Florida]↗

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce↗

Abridged spectral matrix inversion: parametric fitting of X-ray fluorescence spectra following integrative data reduction

Recent improvements in both X-ray detectors and readout speeds have led to a substantial increase in the volume of X-ray fluorescence data being produced at synchrotron facilities. This in turn results in increased challenges associated with processing and fitting such data, both temporally and computationally. Herein an abridging approach is described that both reduces and partially integrates X-ray fluorescence (XRF) data sets to obtain a fivefold total improvement in processing time with negligible decrease in quality of fitting. The approach is demonstrated using linear least-squares matrix inversion on XRF data with strongly overlapping fluorescent peaks. This approach is applicable to any type of linear algebra based fitting algorithm to fit spectra containing overlapping signals wherein the spectra also contain unimportant (non-characteristic) regions which add little (or no) weight to fitted values, e.g. energy regions in XRF spectra that contain little or no peak information.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Feature Analysis, Tracking, and Data Reduction: An Application to Multiphase Reactor Simulation MFiX-Exa for In-Situ Use Case

As we enter the exascale computing regime, powerful supercomputers continue to produce much higher amounts of data than what can be stored for offline data processing. To utilize such high compute capabilities on these machines, much of the data processing needs to happen in situ, when the full high-resolution data is available at the supercomputer memory. In this article, we discuss our MFiX-Exa simulation, which models multiphase flow by tracking a very large number of particles through the simulation domain. In one of the use cases, the carbon particles interact with air to produce carbon dioxide bubbles from the reactor. These bubbles are of primary interest to the domain experts for these simulations. For this particle-based simulation, we propose a streaming technique that can be deployed in situ to efficiently identify the bubbles, track them over time, and use them to down-sample the data with minimal loss in these features.

97 MATHEMATICS AND COMPUTING↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

An Insight-Centric Paradigm for Data Reduction and Inference Speed Improvement at the Scurry Area Canyon Reef Operator’s Committee (SACROC) Unit

The poster presents the work conducted under SMART focusing on using insight-centric approach to design a meaningful proxy for machine learning. Domain insights are critical not just in understanding the prediction results but also in designing the model. This study demonstrated that a single meaningful scaler (as an extreme case) can effectively replace full-size 3D geologic properties. The model's accuracies are on par with other models, and it is the fastest model to predict all test cases, 5,000 times faster than traditional simulations.

Shih, Chung Yan↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗