Precision Cascade: A novel algorithm for multi-precision extreme compression
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Lossless compressors have very low compression ratios that do not meet the needs of today's large-scale scientific applications that produce vast volumes of data. Error-bounded lossy compression (EBLC) is considered a critical technique for the success of scientific research. Although EBLC allows users to set an error bound for the compression, users have been unable to specify the requirements on the compression quality, limiting practical use. Our contributions are: (1) We formulate the problem of configuring EBLC to preserve a user-defined metric as an optimization problem. This allows many classes of new metrics to be preserved, which improves over current practices. (2) We present a framework, OptZConfig, that can adapt to improvements in the search algorithm, compressor, and metrics with minimal changes, enabling future advancements in this area. (3) We demonstrate the advantages of our approach against the leading methods to configure compressors to preserve specific metrics. Here, our approach improves compression ratios against a specialized compressor by up to 3 x, has a 56x speedup over FRaZ, 1000x speedup over MGARD-QOI post tuning, and 110x speedup over systematic approaches which had not been bounded by compressors before.
Today's scientific simulations are producing vast volumes of data that cannot be stored and transferred efficiently because of limited storage capacity, parallel I/O bandwidth, and network bandwidth. The situation is getting worse over time because of the ever-increasing gap between relatively slow data transfer speed and fast-growing computation power in modern supercomputers. Error-bounded lossy compression is becoming one of the most critical techniques for resolving the big scientific data issue, in that it can significantly reduce the scientific data volume while guaranteeing that the reconstructed data is valid for users because of its compression-error-bounding feature. In this paper, we present a novel error-bounded lossy compressor based on a state-of-the-art prediction-based compression framework. Our solution exhibits substantially better compression quality than all of the existing error-bounded lossy compressors, with comparable compression speed. Specifically, our contribution is threefold. (1) We provide an in-depth analysis of why the best-existing prediction-based lossy compressor can only minimally improve the compression quality. (2) We propose a dynamic spline interpolation approach with a series of optimization strategies that can significantly improve the data prediction accuracy, substantially improving the compression quality in turn. (3) We perform a thorough evaluation using six real-world scientific simulation datasets across different science domains to evaluate our solution vs. all other related works. Experiments show that the compression ratio of our solution is higher than that of the second-best lossy compressor by 20%similar to 460% with the same error bound in most of the cases.
As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.
Checkpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications.
OBJECTIVES: The purpose of this study was to evaluate the lossy Joint Photographic Experts Group compression for endodontic pretreatment digital radiographs. STUDY DESIGN: Fifty clinical charge-coupled device-based, digital radiographs depicting periapical areas were selected. Each image was compressed at 2, 4, 8, 16, 32, 48, and 64 compression ratios. One root per image was marked for examination. Images were randomized and viewed by four clinical observers under standardized viewing conditions. Each observer read the image set three times, with at least two weeks between each reading. Three pre-selected sites per image (mesial, distal, apical) were scored on a five-scale score confidence scale. A panel of three examiners scored the uncompressed images, with a consensus score for each site. The consensus score was used as the baseline for assessing the impact of lossy compression on the diagnostic values of images. The mean absolute error between consensus and observer scores was computed for each observer, site, and reading session. RESULTS: Balanced one-way analysis of variance for all observers indicated that for compression ratios 48 and 64, there was significant difference between mean absolute error of uncompressed and compressed images (P <.05). After converting the five-scale score to two-level diagnostic values, the diagnostic accuracy was strongly correlated (R (2) = 0.91) with the compression ratio. CONCLUSION: The results of this study suggest that high compression ratios can have a severe impact on the diagnostic quality of the digital radiographs for detection of periapical lesions.
Simulation based scientific applications generate increasingly large amounts of data on high-performance computing (HPC) systems. To allow data to be stored and analyzed efficiently, data compression is often utilized to reduce the volume and velocity of data. However, a question often raised by domain scientists is the level of compression that can be expected so that they can make more informed decisions, balancing between accuracy and performance. In this letter, we propose a deep neural network based approach for estimating the compressibility of scientific data. To train the neural network, we build both general features as well as compressor-specific features so that the characteristics of both data and lossy compressors are captured in training. Our approach is demonstrated to outperform a prior analytical model as well as a sampling based approach in the case of a biased estimation, i.e., for SZ. However, for the unbiased estimation (i.e., ZFP), the sampling based approach yields the best accuracy, despite the high overhead involved in sampling the target dataset.
Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. Here, to bridge this gap, we propose FedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show that FedEFsz improves the fairness across different benchmarks by up to 60.88% and meanwhile reduces the communication traffic by up to 315×.
Lossy compression techniques have demonstrated promising results in significantly reducing the scientific data size while guaranteeing the compression error bounds. However, one important yet often neglected side effect of lossy scientific data compression is its impact on the performance of parallel I/O. Our key observation is that the compressed data size is often highly skewed across processes in lossy scientific compression. To understand this behavior, we conduct extensive experiments where we apply three lossy compressors MGARD, ZFP, and SZ, which are specifically designed and optimized for scientific data, to three real-world scientific applications Gray-Scott simulation, WarpX, and XGC. Our analysis result demonstrates that the size of the compressed data is always skewed even if the original data is evenly decomposed among processes. Such skewness widely exists in different scientific applications using different compressors as long as the information density of the data varies across processes. We then systematically study how this side effect of lossy scientific data compression impacts the performance of parallel I/O. We observe that the skewness in the sizes of the compressed data often leads to I/O imbalance, which can significantly reduce the efficiency of I/O bandwidth utilization if not properly handled. In addition, writing data concurrently to a single shared file through MPI-IO library is more sensitive to the unbalanced I/O loads. Therefore, we believe our research community should pay more attention to the unbalanced parallel I/O caused by lossy scientific data compression.
Distribution system state estimation requires reliable aggregation of the measured data. However, the large volume of the measured data imposes a significant stress on the underlying communication infrastructure. With the challenges associated with measurement availability, current distribution systems are typically unobservable. To cope with the unobservability issue, compressive sensing theory allows us to recover system state information from a small number of measurements provided the states of the distribution system exhibit sparsity. In this paper, we evaluate the robustness of an updated Kalman filtered modified compressive sensing (KF-ModCS) technique that dynamically estimates the grid states using a small fraction of measured data. In practice, measurements used for sparsity based state estimation may also be intermittent due to communication network induced losses. Further, to understand the effect of packet losses on KF-ModCS, we provide an upper bound for the expected variances of the state estimation error for a given rate of information loss. This upper bound is further improved if the support set of the sparse signal that characterizes the state dynamics does not change over time and/or the reduced model is observable. Simulations based on two practical data sets collected from actual customers in a distribution grid validate the theoretical results.
As the growth of data sizes continues to outpace computational resources, there is a pressing need for data reduction techniques that can significantly reduce the amount of data and quantify the error incurred in compression. Compressing scientific data presents many challenges for reduction techniques since it is often on non-uniform or unstructured meshes, is from a high-dimensional space, and has many Quantities of Interests (QoIs) that need to be preserved. To illustrate these challenges, we focus on data from a large scale fusion code, XGC. XGC uses a Particle-In-Cell (PIC) technique which generates hundreds of PetaBytes (PBs) of data a day, from thousands of timesteps. XGC uses an unstructured mesh, and needs to compute many QoIs from the raw data, f.One critical aspect of the reduction is that we need to ensure that QoIs derived from the data (density, temperature, flux surface averaged momentums, etc.) maintain a relative high accuracy. We show that by compressing XGC data on the high-dimensional, nonuniform grid on which the data is defined, and adaptively quantizing the decomposed coefficients based on the characteristics of the QoIs, the compression ratios at various error tolerances obtained using a multilevel compressor (MGARD) increases more than ten times. We then present how to mathematically guarantee that the accuracy of the QoIs computed from the reduced f is preserved during the compression. We show that the error in the XGC density can be kept under a user-specified tolerance over 1000 timesteps of simulation using the mathematical QoI error control theory of MGARD, whereas traditional error control on the data to be reduced does not guarantee the accuracy of the QoIs.
As the number of operational Phasor Measurement Units (PMUs) is increasing, compression is becoming crucial for efficiently transferring and storing enormous amount of synchrophasor data continuously measured from a power system. Some utilities in the U.S. have already adopted a nonlinear compression technique, called Swinging Door Algorithm (SDA), assuming compressed data meets technical and regulatory requirements without resulting in any degradation of the measurements. This paper studies impacts of this compression algorithm on field measured and simulated PMU data, and presents results from spectral analyses of raw and decompressed measurements. The results show blind application of the SDA compression alters frequency content in the measurements and introduces false harmonics and aliasing. These can lead to false conclusions being drawn from the measurements about the power system condition.
Large atomic-orbital (AO) basis sets of at least triple and preferably quadruple-ζ (QZ) size are required to adequately converge Kohn–Sham density functional theory (DFT) calculations toward the complete basis set limit. However, incrementing the cardinal number by one nearly doubles the AO basis dimension, and the computational cost scales as the cube of the AO dimension, so this is very computationally demanding. Here, in this work, we develop and test a threshold-based natural atomic orbital (NAO) scheme in which ϵ-NAOs are obtained as eigenfunctions of atomic blocks of the density matrix in a one-center orthogonalized representation. This enables compression of the AO basis that is optimal for a given threshold, 10 –ϵ , by discarding NAOs with occupation numbers below that threshold. Extensive pilot test calculations using the Hartree–Fock functional and taking the converged density matrix as input suggest that a threshold of 10 –5 can yield a compression factor (ratio of AO to compressed ϵ-NAO dimension) between 2.5 and 4.5 for the QZ pc-3 basis. The errors in relative energies are typically less than 0.1 kcal/mol when the compressed basis is used instead of the uncompressed basis. Between 10 and 100 times smaller errors (i.e., usually less than 0.01 kcal/mol) can be obtained with a threshold 10 –7 , while the compression factor is typically between 2 and 2.5.
Increases in X-ray brightness from synchrotron light sources lead to a requirement for higher frame rates from hybrid pixel array detectors (HPADs), while also favoring charge integration over photon counting. However, transfer of the full uncompressed data will begin to constrain detector design, as well as limit the achievable continuous frame rate. Here a data compression scheme that is easy to implement in a HPAD's application-specific integrated circuit (ASIC) is described, and how different degrees of compression affect image quality in ptychography, a commonly employed coherent imaging method, is examined. Using adaptive encoding quantization, it is shown in simulations that one can digitize signals up to 16383 photons per pixel (corresponding to 14 bits of information) using only 8 or 9 bits for data transfer, with negligible effect on the reconstructed image.
Explore the source record for details and available documents.
Not provided.
Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.
Explore the source record for details and available documents.