Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compressing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Strategies for on-chip digital data compression for X-ray pixel detectors

Here, the continued desire for X-ray pixel detectors with higher frame rates will stress the ability of application-specific integrated circuit (ASIC) designers to provide sufficient off-chip bandwidth to reach continuous frame rates in the 1 MHz regime. To move from the current 10 kHz to the 1 MHz frame rate regime, ASIC designers will continue to pack as many power-hungry high-speed transceivers at the periphery of the ASIC as possible. In this paper, however, we present new strategies to make the most efficient use of the off-chip bandwidth by utilizing data compression schemes for X-ray photon-counting and charge-integrating pixel detectors. In particular, we describe a novel in-pixel compression scheme that converts from analog to digital converter units to encoded photon counts near the photon Poisson noise level and achieves a compression ratio of >1.5x independent of the dataset. In addition, we describe a simple yet efficient zero-suppression compression scheme called "zeromask" (ZM) located at the ASIC's edge before streaming data off the ASIC chip. ZM achieves average compression ratios of >4x, >7x, and >8x for high-energy X-ray diffraction, ptychography, and X-ray photon correlation spectroscopy datasets, respectively. We present the conceptual designs, register-transfer level block diagrams, and the physical ASIC implementation of these compression schemes in 65 nm CMOS. When combined, these two digital compression schemes could increase the effective off-chip bandwidth by a factor of 6-12x.

47 OTHER INSTRUMENTATION↗

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3↗

Transported PDF Modeling of Compressible Turbulent Reactive Flows by using the Eulerian Monte Carlo Fields Method

Although the transported probability density function (PDF) method has been developed for decades, its application has been mainly focused on the low-Mach number flow problems. This work extends the transported PDF method to compressible flow problems. The Eulerian Monte Carlo fields (EMCF) solution method is employed to solve the transported PDF equation for compressible flow problems. A pseudo stagnation enthalpy is introduced and its stochastic partial differential equation is derived to ensure total energy conservation numerically. A new mixing model called interaction by partial exchange with mean (IPEM) is introduced to expand the available choices of mixing models for the EMCF method. The consistency of the EMCF method is examined for solving the transported PDF equation. Numerical implementation details are discussed, such as the density coupling between the compressible flow solver and the EMCF solver, discretization schemes for the mixing terms and the stochastic terms. The implemented compressible flow solver coupled with the EMCF solver is verified and validated in a series of test cases with increasing level of complexity, ranging from a statistically one-dimensional turbulent mixing layer to a self-excited resonance model rocket combustor. It is observed that in general with the increase of compressibility, there is an increase in the sensitivity of the modeling results to the different models and algorithms. This makes it necessary to develop a thorough understanding of the model sensitivity in order to develop a robust and accurate simulation solver for highly compressible turbulent reactive flows. The thermo-acoustic instability inside the model rocket combustor case is captured reasonably, which demonstrates the overall capability of the developed compressible turbulent combustion solver based on the transported PDF method.

42 ENGINEERING↗

Sound speed measurements in shock-compressed cemented Tungsten carbides at pressures up to 100 GPa

To gain insights into thermodynamic states attained during shock compression of cemented tungsten carbide with 3.7 wt.% cobalt binder, we present results of longitudinal sound (release wave) speed measurements and their analysis at peak stresses up to 100 GPa (volumetric compression ratio ~ 15%). The sound speeds are determined using front-surface impact and release-wave overtake plate impact experimental configurations using laser interferometry. The measured sound speed data along with estimates for bulk sound speeds obtained using the fourth-order Birch-Murnaghan EoS and thermodynamics are used to determine the longitudinal moduli and shear moduli of shocked tungsten carbide at the various peak compression states attained in the experiments. Here, the longitudinal sound speeds were found to increase linearly with volume compression ratio from 6.97 ± 0.010 km/s at ambient conditions to 8.26 ± 0.156 km/s at a volume compression ratio of ~ 15%. The corresponding longitudinal elastic moduli also increase nearly linearly with the volume compression ratio but remain consistently lower than their theoretical predictions based on continuum models with no damage. Also, the sensitivity of shear moduli to pressure, as predicted by the Steinberg-Guinan model, is reduced substantially and the shear moduli of cemented WC with 3.7 wt.% Co remains nearly constant at ~ 310 GPa at the various peak compression stress states investigated in the present study.

36 MATERIALS SCIENCE↗

A 28 nm multiply-accumulate ASIC architecture for on-chip data compression in MHz frame rate X-ray and electron pixel detectors

Modern X-ray detector systems urgently require compact, efficient, and fast data compression schemes to handle the transmission of big data from pixel arrays, enabling frame rates in the MHz regime. Here, in this work, a data compression ASIC that implements a streaming fixed-length lossy compression scheme is introduced and analyzed, proving the feasibility and benefits of on-chip compression. The compression scheme utilizes a vector matrix product logic, which performs a number of floating-point multiplications, additions, and accumulations. The logic is verified, synthesized, and shown to fit in the area resource available for the X-ray detector under study, which comprises 192 × 168 pixels each of 12-bit width, and having a total area of 20 mm× 20 mm, about 2 mm× 20 mm of which are available for the digital logic. Several system architectures, precisions, and compression ratios ranging from 100 to 250 were analyzed to pave the way for on-chip fixed-length compression (e.g., principal component analysis, singular value decomposition) and data reduction (e.g., azimuthal integration) for X-ray and electron detectors.

Data compression↗

zPerf: A Statistical Gray-Box Approach to Performance Modeling and Extrapolation for Scientific Lossy Compression

With the scaling up of simulation-based scientific discovery on high-performance computing systems, the disparity between compute and I/O has increased, forcing domain scientists to save only a small amount of simulation data to persistent storage. This can result in the loss of essential physics fields that are needed for data analysis. While error-bounded lossy compression has made tremendous progress in bridging the gap between compute and I/O, the lack of understanding of compression performance remains a key hurdle to its wide adoption. Here, in this work, we present zPerf, a statistical gray-box performance modeling approach for scientific lossy compression. Our contributions are threefold: 1) We develop zPerf to estimate the performance of lossy compression techniques, based on in-depth understanding and statistical modeling for data features and core compression metrics; 2) We demonstrate the in-detailed implementation of zPerf using two case studies, where we derive the performance modeling for SZ and ZFP, two leading lossy compressors; 3) We evaluate the effectiveness of zPerf on real-world datasets across various domains. Based on the evaluation, we demonstrate the efficacy of the zPerf performance model; 4) We further discuss three case studies where zPerf is applied to extrapolate the compression ratio of SZ and ZFP with alternative encoding schemes as well as ZFP with an alternative transform scheme. Through the case studies, we demonstrate the potential of zPerf for exploring the design space of lossy compression, which has hardly been studied in the literature.

97 MATHEMATICS AND COMPUTING↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

Stability Analysis of Inline ZFP Compression for Floating-Point Data in Iterative Methods

Currently, the dominating constraint in many high performance computing applications is data capacity and bandwidth, in both internode communications and even moreso in intranode data motion. A new approach to address this limitation is to make use of data compression in the form of a compressed data array. Storing data in a compressed data array and converting to standard IEEE-754 types as needed during a computation can reduce the pressure on bandwidth and storage. However, repeated conversions (lossy compression and decompression) introduce additional approximation errors, which need to be shown to not significantly affect the simulation results. Here, we extend recent work that analyzed the error of a single use of compression and decompression of the ZFP compressed data array representation to the case of time-stepping and iterative schemes, where an advancement operator is repeatedly applied in addition to the conversions. We show that the accumulated error for iterative methods involving fixed-point and time evolving iterations is bounded under standard constraints. An upper bound is established on the number of additional iterations required for the convergence of stationary fixed-point iterations. An additional analysis of traditional forward and backward error of stationary iterative methods using ZFP compressed arrays is also presented. The results of several 1D, 2D, and 3D test problems are provided to demonstrate the correctness of the theoretical bounds.

97 MATHEMATICS AND COMPUTING↗

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression↗

Phase Transitions in (Mg,Fe) 2 SiO 4 Olivine under Shock Compression

The response of forsterite, Mg 2 SiO 4 , under dynamic compression is of fundamental importance for understanding its phase transformations and high-pressure behavior. Here, we have carried out an in situ X-ray diffraction study of laser-shocked polycrystalline and single-crystal forsterite from 19 to 122 GPa using the Matter in Extreme Conditions end-station of the Linac Coherent Light Source. Under laser-based shock loading, forsterite does not transform to the high-pressure equilibrium assemblage of MgSiO 3 bridgmanite and MgO periclase, as has been suggested previously. Instead, we observe forsterite and forsterite III, a metastable polymorph of Mg 2 SiO 4 , coexisting in a mixed-phase region from 33 to 75 GPa for both polycrystalline and single-crystal samples. Densities inferred from X-ray diffraction are consistent with earlier gas-gun shock data. At higher stress, the response is sample-dependent. Polycrystalline samples undergo amorphization above 79 GPa. For [010]- and [001]-oriented crystals, a mixture of crystalline and amorphous material is observed to 108 GPa, whereas the [100]-oriented forsterite adopts an unknown phase at 122 GPa. The first two sharp diffraction peaks of amorphous Mg 2 SiO 4 show a similar trend with compression as those observed for MgSiO 3 in both recent static- and laser-driven shock experiments. This study provides new insight into the transformation of forsterite under nanosecond-duration shock loading. This work emphasizes the importance of formation of metastable phases along the Hugoniot and adds to evidence that the 300-K single-crystal diamond anvil cell experiments have relevance for understanding structures formed under shock compression. In particular, the metastable phase forsterite III has now been shown to form under dynamic compression from 10s to 100s of nanoseconds as well as under 300-K static compression. Upon compression to higher pressures, Mg 2 SiO 4 transforms to an amorphous phase. These results have broad relevance for understanding the behavior of silicates under dynamic compression.

36 MATERIALS SCIENCE↗

Effect of Part Size, Displacement Rate, and Aging on Compressive Properties of Elastomeric Parts of Different Unit Cell Topologies Formed by Vat Photopolymerization Additive Manufacturing

Due to its ability to achieve geometric complexity at high resolution and low length scales, additive manufacturing (AM) has increasingly been used for fabricating cellular structures (e.g., foams and lattices) for a variety of applications. Specifically, elastomeric cellular structures offer tunability of compliance as well as energy absorption and dissipation characteristics. However, there are limited data available on compression properties for printed elastomeric cellular structures of different designs and testing parameters. In this work, the authors evaluate how unit cell topology, part size, the rate of compression, and aging affect the compressive response of polyurethane-based simple cubic, body-centered, and gyroid structures formed by vat photopolymerization AM. Finite element simulations incorporating hyperelastic and viscoelastic models were used to describe the data, and the simulated results compared well with the experimental data. Of the designs tested, only the parts with the body-centered unit cell exhibited differences in stress–strain responses at different part sizes. Of the compression rates tested, the highest displacement rate (1000 mm/min) often caused stiffer compressive behavior, indicating deviation from the quasi-static assumption and approaching the intermediate rate response. The cellular structures did not change in compression properties across five weeks of aging time, which is desirable for cushioning applications. This work advances knowledge on the structure–property relationships of printed elastomeric cellular materials, which will enable more predictable compressive properties that can be traced to specific unit cell designs.

36 MATERIALS SCIENCE↗

Lithium plating induced degradation during fast charging of batteries subjected to compressive loading

Here we report the lithium plating associated capacity loss during fast charging of compressively loaded lithium-ion batteries (LIBs). The charging and discharging of LIB under compressive loading during service may affect the cell performance or initiate localized defects in the electrodes. Pouch cells of capacity 20 mAh were compressively loaded to nominal pressures of 0–440 kPa and subjected to 10 cycles of fast charging at 1 C and 4 C. Experimental results show that cells charged at 4 C-rate experienced significant capacity fade, and applying compressive loads exacerbated the capacity loss. The coulombic efficiency study shows that active lithium loss was higher for the initial cycles before gradually reducing to a minimal capacity loss for the tenth charging cycle. The cell voltage relaxation immediately after charging was monitored to identify the stripping of plated lithium after fast charging cycles and showed that the duration of lithium stripping was higher for cells under mechanical compressive loading. Scanning electron microscopy (SEM) and electron paramagnetic resonance spectroscopy (EPR) characterization of the anode showed significantly higher lithium deposits on the anodes charged at a 4 C rate under compressive loads. These results indicate that applied mechanical compression causes increased lithium plating during fast charging of batteries.

25 ENERGY STORAGE↗

Dynamic compression effects of H 2 ⁡O in a dynamic diamond anvil cell: Origin of metastable ice VII and its crystal growth kinetics

We report on the structural verification of metastable ice VII solidifying in the phase space of ice VI at 1.80 GPa at room temperature. Using time-resolved (TR) x-ray diffraction and TR ruby luminescence paired with high-speed microphotography utilizing a dynamic diamond anvil cell, an initial compression rate range from 0.12 to 95.84 GPa/s was explored. The solidification pressure of metastable ice VII has a potential sigmoidal dependence upon compression rate with a turnover compression rate of ∼80 GPa/s. The preferred crystallization of ice VII in the stability field of ice VI is due to the increased nucleation rate of ice VII over ice VI at 1.77 GPa that is driven by the surface energy difference between the liquid and solid phases along with the change in Gibbs free energy of solidification. The dynamic pressure-volume–compression behaviors of ice phases (VI and VII) show a lattice stiffening in both phases, especially during the compression loading. It is also found that the compression rate greatly affects the solid-solid phase transition between ice VI and VII but does not affect the liquid-solid transition between water and ice VI as much. Lastly, a third phase transition was found to occur after metastable ice VII transforms into high-density amorphous (HDA) ice, which could be a disordered hydrogen-bonded network configuration of ice VII forming out of HDA ice facilitated by the decoupling of the oxygen movement and reorientation of the H 2⁡ O molecule. These results demonstrate the complexity of a seemingly simple molecule H 2⁡ O, how it can readily change its static properties with the modification of (de)compression rate, and highlight the need to use multiple TR structural and spectroscopic probes at higher time resolutions to realize the most comprehensive understanding.

Chemical bonding↗

Ion versus Electron Heating in Compressively Driven Astrophysical Gyrokinetic Turbulence

The partition of irreversible heating between ions and electrons in compressively driven (but subsonic) collisionless turbulence is investigated by means of nonlinear hybrid gyrokinetic simulations. We derive a prescription for the ion-to-electron heating ratio Q i /Q e as a function of the compressive-to-Alfvénic driving power ratio P compr /P AW , of the ratio of ion thermal pressure to magnetic pressure β i , and of the ratio of ion-to-electron background temperatures T i /T e . It is shown that Q i /Q e is an increasing function of P compr /P AW . When the compressive driving is sufficiently large, Q i /Q e approaches ≃P compr /P AW . This indicates that, in turbulence with large compressive fluctuations, the partition of heating is decided at the injection scales, rather than at kinetic scales. Analysis of phase-space spectra shows that the energy transfer from inertial-range compressive fluctuations to sub-Larmor-scale kinetic Alfvén waves is absent for both low and high β i , meaning that the compressive driving is directly connected to the ion-entropy fluctuations, which are converted into ion thermal energy. This result suggests that preferential electron heating is a very special case requiring low βi and no, or weak, compressive driving. Our heating prescription has wide-ranging applications, including to the solar wind and to hot accretion disks such as M87 and Sgr A*.

79 ASTRONOMY AND ASTROPHYSICS↗

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

Black-box statistical prediction of lossy compression ratios for scientific data

Lossy compressors are increasingly adopted in scientific research, tackling volumes of data from experiments or parallel numerical simulations and facilitating data storage and movement. In contrast with the notion of entropy in lossless compression, no theoretical or data-based quantification of lossy compressibility exists for scientific data. Users rely on trial and error to assess lossy compression performance. As a strong data-driven effort toward quantifying lossy compressibility of scientific datasets, we provide a statistical framework to predict compression ratios of lossy compressors. Our method is a two-step framework where (i) compressor-agnostic predictors are computed and (ii) statistical prediction models relying on these predictors are trained on observed compression ratios. Proposed predictors exploit spatial correlations and notions of entropy and lossyness via the quantized entropy. We study 8+ compressors on 6 scientific datasets and achieve a median percentage prediction error less than 12%, which is substantially smaller than that of other methods while achieving at least a 8.8× speedup for searching for a specific compression ratio and 7.8× speedup for determining the best compressor out of a collection.

97 MATHEMATICS AND COMPUTING↗

Chapter 22: Compressed Air Evaluation Protocol. The Uniform Methods Project: Methods for Determining Energy Efficiency Savings for Specific Measures (September 2011 - August 2020)

Compressed-air systems are used widely throughout industry for many operations, including pneumatic tools, packaging and automation equipment, conveyors, and other industrial process operations. Compressed-air systems are defined as a group of subsystems composed of air compressors, air treatment equipment, controls, piping, pneumatic tools, pneumatically powered machinery, and process applications using compressed air. A compressed-air system has three primary functional subsystems: supply, distribution, and demand. Air compressors are the primary energy consumers in a compressed-air system and are the primary focus of this protocol. The two compressed-air energy efficiency measures specifically addressed in this protocol are: high-efficiency/variable speed drive (VSD) compressor replacing modulating, load/unload, or constant-speed compressor; compressed-air leak survey and repairs. This protocol provides direction on how to reliably verify savings from these two measures using a consistent approach for each.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗