Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool for compressing large data matrices that arise in simulations of partial differential equations (PDEs), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. Here, to address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices in situ. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

Column subset selection↗

Efficient data reduction for time-of-flight neutron scattering experiments on single crystals

Event-mode data collection presents remarkable new opportunities for time-of-flight neutron scattering studies of collective excitations, diffuse scattering from short-range atomic and magnetic structures, and neutron crystallography. In these experiments, large volumes of the reciprocal space are surveyed, often using different wavelengths and counting times. These data then have to be added together, with accurate propagation of the counting errors. This paper presents a statistically correct way of adding and histogramming the data for single-crystal time-of-flight neutron scattering measurements. In order to gain a broader community acceptance, particular attention is given to improving the efficiency of calculations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybrid Approaches for Data Reduction of Spatiotemporal Scientific Applications

Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions.

Li, Xiao↗

Machine Learning Techniques for Data Reduction of Climate Applications

Scientists conduct large-scale simulations to compute derived quantities-of-interest (QoI) from primary data. Often, QoI are linked to specific features, regions, or time intervals, such that data can be adaptively reduced without compromising the integrity of QoI. For many spatiotemporal applications, these QoI are binary in nature and represent presence or absence of a physical phenomenon. We present a pipelined compression approach that first uses neural-network-based techniques to derive regions where QoI are highly likely to be present. Then, we employ a Guaranteed Autoencoder (GAE) to compress data with differential error bounds. GAE uses QoI information to apply low-error compression to only these regions. This results in overall high compression ratios while still achieving downstream goals of simulation or data collections. Experimental results are presented for climate data generated from the E3SM Simulation model for downstream quantities such as tropical cyclone and atmospheric river detection and tracking. These results show that our approach is superior to comparable methods in the literature.

Li, Xiao [University of Florida]↗

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce↗

The Exploitation of Data Reduction for Visualization

The disparity between the computational speed and storage bandwidth, as demonstrated in Figure 1, is a well known problem that grows with each successive generation. The visualization community is principally responding to this issue by using in situ to reduce which data must be written to storage. However, other communities are taking different, possibly complementary approaches. In particular, data compression is a common general approach to reduce storage demands. Data compression technologies are typically not designed with post processing in mind. The principal metrics measured are compression ratio, the improved bandwidth to storage, and the error introduced. It is assumed that data is inflated to its full size before any post processing can happen. Although when talking about bandwidth disparities, HPC’s dirty little secret is that no part of the memory nor interconnect hardware is increasing at the rate of computation. For example, the Summit supercomputer has a peak computation rate almost 10 times its predecessor, Titan, but only about 4 times the memory, less than twice the aggregate memory bandwidth, and almost no improvement in the interconnect bisection bandwidth. Naively inflating data for post processing does not help with limitations in the memory and interconnect systems.

97 MATHEMATICS AND COMPUTING↗

An Insight-Centric Paradigm for Data Reduction and Inference Speed Improvement at the Scurry Area Canyon Reef Operator’s Committee (SACROC) Unit

The poster presents the work conducted under SMART focusing on using insight-centric approach to design a meaningful proxy for machine learning. Domain insights are critical not just in understanding the prediction results but also in designing the model. This study demonstrated that a single meaningful scaler (as an extreme case) can effectively replace full-size 3D geologic properties. The model's accuracies are on par with other models, and it is the fastest model to predict all test cases, 5,000 times faster than traditional simulations.

Shih, Chung Yan↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

AMM: Adaptive Multilinear Meshes

Adaptive representations are increasingly indispensable for reducing the in-memory and on-disk footprints of large-scale data. Usual solutions are designed broadly along two themes: reducing data precision, e.g., through compression, or adapting data resolution, e.g., using spatial hierarchies. Additionally, recent research suggests that combining the two approaches, i.e., adapting both resolution and precision simultaneously, can offer significant gains over using them individually. However, there currently exist no practical solutions to creating and evaluating such representations at scale. In this work, we present a new resolution-precision-adaptive representation to support hybrid data reduction schemes and offer an interface to existing tools and algorithms. Through novelties in spatial hierarchy, our representation, Adaptive Multilinear Meshes (AMM), provides considerable reduction in the mesh size. AMM creates a piecewise multilinear representation of uniformly sampled scalar data and can selectively relax or enforce constraints on conformity, continuity, and coverage, delivering a flexible adaptive representation. AMM also supports representing the function using mixed-precision values to further the achievable gains in data reduction. We describe a practical approach to creating AMM incrementally using arbitrary orderings of data and demonstrate AMM on six types of resolution and precision datastreams. By interfacing with state-of-the-art rendering tools through VTK, we demonstrate the practical and computational advantages of our representation for visualization techniques. With an open-source release of our tool to create AMM, we make such evaluation of data reduction accessible to the community, which we hope will foster new opportunities and future data reduction schemes.

97 MATHEMATICS AND COMPUTING↗

Machine-learning-assisted automation of single-crystal neutron diffraction

Neutron scattering is a powerful but expensive technique to study materials and discover new matter. Advanced detector technology has significantly improved the efficiency of neutron experiments, increasing the complexity of neutron data reduction and analysis. Machine learning (ML) brings new directions for neutron diffraction data reduction and experiment operation. Here, this work presents an ML-assisted data reduction and analysis method for precise recognition of Bragg peaks and the corresponding regions of interest; it can then automatically screen and align a measured crystal using the recognized peaks, and subsequently plan and optimize the data collection with user-provided information and uncertainty quantification values of detected peaks. This method shows robust performance in different complex sample environments and enables automated single-crystal neutron diffraction.

47 OTHER INSTRUMENTATION↗

Robustness of the smartpixels classifier for different simulated sensor geometries and non-ideal detector conditions

Pixel tracking detectors at upcoming collider experiments will see unprecedented charged-particle densities. Real-time data reduction on the detector will enable higher granularity and faster readout, possibly enabling the use of the pixel detector in high-rate online event selection, such as the ATLAS or CMS first-level trigger systems. This data reduction can be accomplished with a neural network (NN) in the readout chip bonded with the sensor that recognizes and rejects tracks with low transverse momentum (p T ) based on the geometrical shape of the charge deposition (“cluster”). To design viable detectors for deployment, the dependence of the NN as a function of the sensor geometry, external magnetic field, irradiation, and noise must be understood. In this paper, we present first studies of the efficiency and data reduction for planar pixel sensors exploring these parameters. For the CMS HL-LHC sensor geometry, we obtain a signal efficiency of (91.9 ± 0.7)% and a data reduction of (29.7 ± 1.0)%. A smaller sensor pitch in the bending direction improves the p T discrimination, but a larger pitch can be partially compensated with detector thickness. Any accumulated radiation damage also changes the cluster shape, reducing the signal efficiency compared to the baseline by approximately 30–60% in absolute terms, but nearly all of the performance can be recovered through retraining of the network and updating the weights. Finally, the impact of noise was investigated, and retraining the network on noise-injected datasets was found to maintain performance within 6% of the baseline network trained and evaluated on noiseless data. •ASIC-compatible track-momentum classifier is robust in realistic detector conditions.•About 90% signal efficiency and 30% data reduction per layer for CMS HL-LHC geometry.•Single-layer signal efficiency increases for smaller pixel pitch or thicker sensors.•Performance with noise or after radiation damage mostly recovered by retraining.

Shekar, Danush [Illinois U., Chicago] (ORCID:00000↗

Tools for Visualization and Analysis of Small-Angle Neutron Scattering Data: Descriptions and Examples

A great deal of progress has been made in improving the data reduction experience for the SANS instruments at the SNS and HFIR at ORNL. The existing data reduction toolset, drtsans, makes it possible to integrate data analysis and visualization tools into the data reduction scripts, thereby providing new opportunities for more automated data processing for users of the SNS and HFIR. Here, the first set of tools developed is described with usage examples.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ECON-T and ECON-D: Endcap Concentrator ASICs for the CMS HGCAL

With over 6 million channels, the High Granularity Calorimeter (HGCAL) for the CMS HL-LHC Upgrade presents a unique data challenge. The ECON ASICs provide critical on-detector data reduction for the 40 MHz trigger path (ECON-T) and 750 kHz data acquisition path (ECON-D) of the HGCAL. The ASICs, fabricated in 65 nm CMOS, are rad-tolerant (600 Mrad) with low power consumption (<2.5 mW/channel). This presentation is the first comprehensive description of the ECON designs, first functionality and radiation tests for the ECON-T ASIC, and first results from the full production of 75k ECON-D and ECON-T ASICs.

Bergamin, G. [CERN] (ORCID:0000000285758704)↗