Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compressing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Unbalanced Parallel I/O: An Often-Neglected Side Effect of Lossy Scientific Data Compression

Lossy compression techniques have demonstrated promising results in significantly reducing the scientific data size while guaranteeing the compression error bounds. However, one important yet often neglected side effect of lossy scientific data compression is its impact on the performance of parallel I/O. Our key observation is that the compressed data size is often highly skewed across processes in lossy scientific compression. To understand this behavior, we conduct extensive experiments where we apply three lossy compressors MGARD, ZFP, and SZ, which are specifically designed and optimized for scientific data, to three real-world scientific applications Gray-Scott simulation, WarpX, and XGC. Our analysis result demonstrates that the size of the compressed data is always skewed even if the original data is evenly decomposed among processes. Such skewness widely exists in different scientific applications using different compressors as long as the information density of the data varies across processes. We then systematically study how this side effect of lossy scientific data compression impacts the performance of parallel I/O. We observe that the skewness in the sizes of the compressed data often leads to I/O imbalance, which can significantly reduce the efficiency of I/O bandwidth utilization if not properly handled. In addition, writing data concurrently to a single shared file through MPI-IO library is more sensitive to the unbalanced I/O loads. Therefore, we believe our research community should pay more attention to the unbalanced parallel I/O caused by lossy scientific data compression.

Wang, Xinying↗

Ultrafast Error-bounded Lossy Compression for Scientific Datasets

Today's scientific high-performance computing applications and advanced instruments are producing vast volumes of data across a wide range of domains, which impose a serious burden on data transfer and storage. Error-bounded lossy compression has been developed and widely used in the scientific community because it not only can significantly reduce the data volumes but also can strictly control the data distortion based on the user-specified error bound. Existing lossy compressors, however, cannot offer ultrafast compression speed, which is highly demanded by numerous applications or use cases (such as in-memory compression and online instrument data compression). In this paper, we propose a novel ultrafast error-bounded lossy compressor that can obtain fairly high compression performance on both CPUs and GPUs and with reasonably high compression ratios. The key contributions are threefold. (1) We propose a generic error-bounded lossy compression framework---called SZx---that achieves ultrafast performance through its novel design comprising only lightweight operations such as bitwise and addition/subtraction operations, while still keeping a high compression ratio. (2) We implement SZx on both CPUs and GPUs and optimize the performance according to their architectures. (3) We perform a comprehensive evaluation with six real-world production-level scientific datasets on both CPUs and GPUs. Experiments show that SZx is 2~16x faster than the second-fastest existing error-bounded lossy compressor (either SZ or ZFP) on CPUs and GPUs, with respect to both compression and decompression.

Yu, Xiaodong↗

Homomorphic data compression for real time photon correlation analysis

The construction of highly coherent X-ray sources, combined with next-generation detectors that are larger and faster, has enabled new research opportunities across the scientific landscape. Among the techniques that benefit most from these advancements is X-ray photon correlation spectroscopy (XPCS), where faster acquisition unlocks the ability to study faster dynamics within samples. However, faster acquisition on larger detectors also introduces unprecedented challenges for online data processing and offline data storage. Such challenges are particularly prominent for XPCS, where real time analyses require simultaneous calculation of all the previously acquired data in the time series. We present a homomorphic compression scheme to effectively reduce the computational time and memory space required for XPCS analysis. Leveraging similarities in the mathematical expression between a matrix-based compression algorithm and the correlation calculation, our approach allows direct operation on the compressed data without their decompression. The offline compression scheme extends storage capacity by a factor of 40 while preserving key features in the lossy compressed data. Meanwhile, the online compression scheme reduces the computational time to below 1 ms, enabling real time calculation of the correlation functions at kHz framerate. Our demonstration of a homomorphic compression of scientific data provides an effective solution to the big data challenge at coherent light sources. Beyond the example shown in this work, the framework can be extended to facilitate real-time operations directly on a compressed data stream for other techniques.

36 MATERIALS SCIENCE↗

Real-time and post-hoc compression for data from Distributed Acoustic Sensing

Distributed Acoustic Sensing (DAS) is an emerging sensing technology that records the strain-rate along fiber optic cables at high spatial and temporal resolution. This technique is becoming a popular tool in seismology, hydrology, and other subsurface monitoring applications. However, due to the large coverage (10’s of km) and high density of measurements (1m spacing at 100’s of Hz), a DAS installation could produce terabytes of data records per day. Because many DAS instruments are deployed in remote locations, this large data size poses significant challenges to its transfer and storage. In this paper, we explore lossless compression methods to reduce the storage requirement in both real-time and post-hoc scenarios. Here we propose a two-stage compression method to improve the compression ratio and compression speed. This two-stage compression method could reduce the storage requirement by 40%, which is 20% more than other lossless methods, such as ZSTD. We demonstrate that the compression method could complete its operation well before the DAS instrument needs to output the next file, making it suitable for real-time DAS acquisition. We also implement a parallel compression method for a post-hoc scenario and demonstrate that our method could effectively utilize a parallel computer. With 256 CPU cores, our parallel compression method achieves the speed of 26GB/second.

58 GEOSCIENCES↗

Initial assessment of alternative carbon fiber geometries for design of cost-effective compressive performance: Size effect studies

Carbon fiber provides opportunity to reduce weight in structural composites, including wind turbine blades, due to the material's superior specific stiffness and specific strength compared to alternatives. Despite these advantages, cost and compressive performance are considered weaknesses for carbon fiber products available today. Studies to produce low-cost carbon fiber alternatives, including the use of textile-derived precursor systems, have shown progress and merit through the DOE/ORNL low-cost carbon fiber initiatives. Here, this work focuses on enabling increases in compressive strength through design of the carbon fiber geometry, applicable to both textile and conventional precursor systems, while also providing opportunities to reduce carbon fiber processing costs. Fiber-resin interface and fiber alignment are among the most frequently cited factors controlling composite compressive performance. However, it is believed that there is opportunity in traditionally unexplored routes to increasing compressive strength through alteration of the carbon fiber geometry by increasing the fiber area moment of inertia and/or the fiber perimeter and interfacial area. This paper presents initial results from manufacturing carbon fiber materials to assess the impacts of carbon fiber size on tested composite compressive performance with projected neutral or even beneficial impact on fiber and composite manufacturing economics. Carbon fiber systems with increasing size illustrate a favorable correlation for compressive performance greater than predicted from a micromechanical failure model. The manufacturing and mechanical test results support the hypothesis of this work that alterations to fiber geometry can be used to produce improvements of the compressive strength of carbon fiber reinforced polymers and provide incentive for related work in designing alternative shapes to further enhance compressive performance.

42 ENGINEERING↗

Compression–tension cell with sample manipulator for in situ X‐ray nanotomography experiments

In situ X-ray nanotomography experiments where tensile or compressive force is applied on the sample require specialized equipment. A compression-tension device with fluid flow-through capability has been designed for X-ray nanotomography beamlines. The compression-tension cell is equipped with a triaxial stage for sample alignment and a high sensitivity loadcell for measurement of applied force. To handle the <100 µm samples used for X-ray nanotomography imaging and for loading samples on the compression-tension cell a sample manipulator has been built. The sample manipulator is capable of selecting a single <100 µm particle for nanotomography scanning while viewing multiple samples under an optical microscope. To test the functionality of these two devices an initial compression experiment involving two glass beads was performed. To demonstrate instrument stability two spherical glass beads were compressed from a no load condition until one of the beads fractured. Nanotomography data were collected at each step of increasing compressive force. The experimentally observed contact area of the spherical glass beads was compared with the theoretical estimate using the Hertz analysis. To demonstrate the fluid flow capability, two calcite grains were compressed against each other under a calcite saturated solution. Surface topological changes were observed for the stressed grain contact area.

X-ray tomography↗

cuZ-Checker: A GPU-Based Ultra-Fast Assessment System for Lossy Compressions

Lossy compression is becoming an indispensable technique for the success of today's extreme-scale high-performance computing projects that produce vast volumes of data during scientific simulations or instrument data acquisitions. Comprehensively understanding the compression quality and performance of different lossy compressors is critical to selecting the best-fit compressors and using them properly and efficiently in practice. A few lossy compression assessment tools (e.g., Z-checker) have been developed, but none of them support the execution in a GPU environment. This is a significant gap because many recent extreme-scale applications and lossy compressors (e.g., cuSZ) can run entirely within GPUs. In this work, we develop an efficient lossy compression measuring system (called cuZ-Checker) on the GPU platform, which aims to perform the lossy compression quality and performance assessment completely within the GPU environment. Our contribution is threefold. (1) We develop a novel GPU-based lossy compression measuring framework using a computation pattern-based design approach. This approach classifies the computing-intensive metrics into three categories based on their patterns which creates large opportunities for kernel fusion and data reuse. (2) For each pattern in cuZ-Checker, we develop a CUDA kernel and provide fine-grained optimizations to boost its performance. (3) We thoroughly evaluate our cuZ-checker on a V100 GPU using four real-world scientific application datasets. Experiments show that cuZ-Checker can significantly accelerate the overall lossy compression assessment performance by 23X similar to 31X compared with the OpenMP-based multithreading CPU performance. To the best of our knowledge, this is the first lossy compression measuring system designed for GPU devices.

GPU↗

Taking control of compressible modes: bulk viscosity and the turbulent dynamo

Many polyatomic astrophysical plasmas are compressible and out of chemical and thermal equilibrium, introducing a bulk viscosity into the plasma via the internal degrees of freedom of the molecular composition, directly impacting the decay of compressible modes, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$. This is especially important for small-scale, turbulent dynamo processes in the interstellar medium (ISM), which are known to be sensitive to the effects of compression. To control the viscous properties of $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$, we perform trans-sonic, visco-resistive dynamo simulations with additional bulk viscosity $\nu _{\text{bulk}}$, deriving a new $\nu _{\text{bulk}}$ Reynolds number $\text{Re}_{\text{bulk}}$, and viscous Prandtl number $\text{P}\nu \equiv \text{Re}_{\text{bulk}}/ \text{Re}_{\text{shear}}$, where $\text{Re}_{\text{shear}}$ is the shear viscosity Reynolds number. We derive a framework for decomposing $E_{\rm mag}$ growth rates into incompressible and compressible terms via orthogonal tensor decompositions of $\boldsymbol {\nabla }\otimes \mathrm{{\boldsymbol {\mathit {v}}}}$, where $\mathrm{{\boldsymbol {\mathit {v}}}}$ is the fluid velocity. We find that $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ play a dual role, growing and decaying $E_{\rm mag}$, and that field-line stretching is the main driver of growth, even in compressible dynamos. In the absence of $\nu _{\text{bulk}}$ ($\text{P}\nu \rightarrow \infty$), $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ pile up on small-scales, creating a spectral bottleneck, which disappears for $\text{P}\nu \approx 1$. As $\text{P}\nu$ decreases, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ are dissipated at increasingly larger scales, in turn suppressing incompressible modes through a coupling between high-k modes. We emphasize the importance of further understanding the role of $\nu _{\text{bulk}}$ in compressible astrophysical plasmas, which we estimate could be as strong as the shear viscosity in the cold ISM, and highlight that compressible direct numerical simulations without bulk viscosity have unresolved compressible mode dissipation scales.

MHD↗

Preferential turbulence enhancement in two-dimensional compressions

When initially isotropic three-dimensional (3D) turbulence is compressed along two dimensions, the compression supplies energy directly to the flow components in the compressed directions, while the flow component in the noncompressed direction experiences the effects of compression only indirectly through the nonlinearity of the hydrodynamic equations. Here we present our study of such 2D compressions using numerical simulations. For initially isotropic turbulence, we find that the nonlinearity can be insufficient to maintain isotropy, with the energy components parallel to the compression coming to dominate the turbulent energy, with a number of consequences. Among these are the possibilities for stronger and more easily sustained growth of turbulent energy than in 3D compressions and for an increasing turbulent Mach number even in a compression without thermal losses.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Spatiotemporally Adaptive Compression for Scientific Dataset with Feature Preservation – A Case Study on Simulation Data with Extreme Climate Events Analysis

Scientific discoveries are increasingly constrained by limited storage space and I/O capacities. For time-series simulations and experiments, their data often need to be decimated over timesteps to accommodate storage and I/O limitations. In this paper, we propose a technique that addresses storage costs while improving post-analysis accuracy through spatiotemporal adaptive, error-controlled lossy compression. We investigate the trade-off between data precision and temporal output rates, revealing that reducing data precision and increasing timestep frequency lead to more accurate analysis outcomes. Additionally, we integrate spatiotemporal feature detection with data compression and demonstrate that performing adaptive error-bounded compression in higher dimensional space enables greater compression ratios, leveraging the error propagation theory of a transformation-based compressor. To evaluate our approach, we conduct experiments using the well-known E3SM climate simulation code and apply our method to compress variables used for cyclone tracking. Our results show a significant reduction in storage size while enhancing the quality of cyclone tracking analysis, both quantitatively and qualitatively, in comparison to the prevalent timestep decimation approach. Compared to three state-of-the-art lossy compressors lacking feature preservation capabilities, our adaptive compression framework improves perfectly matched cases in TC tracking by 26.4-51.3% at medium compression ratios and by 77.3-571.1% at large compression ratios, with a merely 5–11% computational overhead.

Gong, Qian↗

Optimizing Error-Bounded Lossy Compression for Scientific Data by Dynamic Spline Interpolation

Today's scientific simulations are producing vast volumes of data that cannot be stored and transferred efficiently because of limited storage capacity, parallel I/O bandwidth, and network bandwidth. The situation is getting worse over time because of the ever-increasing gap between relatively slow data transfer speed and fast-growing computation power in modern supercomputers. Error-bounded lossy compression is becoming one of the most critical techniques for resolving the big scientific data issue, in that it can significantly reduce the scientific data volume while guaranteeing that the reconstructed data is valid for users because of its compression-error-bounding feature. In this paper, we present a novel error-bounded lossy compressor based on a state-of-the-art prediction-based compression framework. Our solution exhibits substantially better compression quality than all of the existing error-bounded lossy compressors, with comparable compression speed. Specifically, our contribution is threefold. (1) We provide an in-depth analysis of why the best-existing prediction-based lossy compressor can only minimally improve the compression quality. (2) We propose a dynamic spline interpolation approach with a series of optimization strategies that can significantly improve the data prediction accuracy, substantially improving the compression quality in turn. (3) We perform a thorough evaluation using six real-world scientific simulation datasets across different science domains to evaluate our solution vs. all other related works. Experiments show that the compression ratio of our solution is higher than that of the second-best lossy compressor by 20%similar to 460% with the same error bound in most of the cases.

Zhao, Kai↗

Accelerating Lossy and Lossless Compression on Emerging BlueField DPU Architectures

Data compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios.

Li, Yuke↗

Real-Time Lossless Compression for Ultra-High-Density Synchrophasor and Point on Wave Data

Modern advanced Phasor Measurement Units (PMUs) are developed with ultra-high reporting rates to meet the demand for monitoring the power systems dynamics in detail. Due to the large volume of data, the communication and storage systems are seriously challenged with the presence of Ultra-High-Density (UHD) synchrophasor and Point on Wave (POW) data. Therefore, it is an urgent task to compress the UHD data for more efficient communication and data storage. This paper proposes several methods to compress the synchrophasor and POW data in a lossless manner. First, an Improved-Time-Series-Special Compression (ITSSC) method is proposed to compress the UHD frequency data. Second, a Delta-difference Huffman method is combined with the TSSC algorithm to compress the UHD phase angle data. Finally, a cyclical high-order delta modulation method is proposed to compress the UHD POW data. The proposed models are extensively tested and compared with different existing lossless compression algorithms using the field-collected synchrophasor and POW data at different reporting rates. The results indicate that the proposed algorithms are efficient in performing lossless compression for the UHD synchrophasor and POW data in real time.

42 ENGINEERING↗

Region-adaptive, Error-controlled Scientific Data Compression using Multilevel Decomposition

The increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest.We conduct the evaluations on two climate use cases – one targeting small-scale, node features and the other focusing on long, areal features. For both use cases, the locations of the features were unknown ahead of the compression. By selecting approximately 16% of the data based on multi-scale spatial variations and compressing those regions with smaller error tolerances than the rest, our approach improves the accuracy of post-analysis by approximately 2 × compared to single-error-bounded compression at the same compression ratio. Using the same error bound for the region of interest, our approach can achieve an increase of more than 50% in overall compression ratio.

Gong, Qian↗

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.

Wang, Daoce [University of Nebraska, Omaha]↗

Determining Compression Characteristics of Honeycomb Material - 19674

The objective is to determine compression test characteristics of stainless steel honeycomb material to be able to represent honeycomb structures accurately in analytical models used to simulate hypothetical accident scenarios of shipping packages. Honeycomb is a material used primary in the aerospace industry due to its high strength-to-weight ratio. Because it doesn't have a shelf life and can withstand high heats, it is an excellent candidate for a structural material in package designs. Honeycomb has orthotropic material properties. The honeycomb currently being evaluated is made of metal ribbons spot welded together to form hexagon pattern between metal plates. The hexagon ribbons are brazed to the metal plates. The complexity of the honeycomb significantly lengthens the simulation time, which is further complicated when the brazing, welding, and imperfections of the material is considered. This means to affectively represent the honeycomb in simulation programs, such as Abaqus, the material needs to be approximated as a uniform orthotropic material. This requires material properties in each of the three directions, T, W, and L. The T direction is defined as the 'strong' direction, perpendicular to the honeycomb sheet. The L direction is parallel the ribbon and the W direction is perpendicular to the ribbon. Honeycomb has 3 different compression stages. Stage 1 is the initial compression where the honeycomb maintains its structural integrity and does not permanently deform from loads from normal operations. Stage 2 is when the initial buckling causes irreversible damage to the honeycomb. Stage 2 is the one we are most interested in because it absorbs the most energy from a hypothetical accident scenario. During Stage 2, the honeycomb fails layer-by-layer, indicating that the more layers, the longer material crushing is sustained, and thus the more energy that is absorbed. Stage 3 is final compression, similar to compressing solid metal. During stage 3 the stress increases with a diminishing rate of elongation and the honeycomb is completely failed where the plates between the honeycomb core sandwich the crushed honeycomb ribbon. The stress-strain graph bellow shows compression test of 3 different samples in the T direction. The red lines divide the three stages. The far left is Stage 1, the middle stage 2, and the right stage 3. The data collected is displayed in the table below. The expressions in the table represent the regression formula that represents each stage of the compression on a stress-strain graph. Stage 1 is assumed to intersect with the origin; data for stage 3 for the crush in W and L directions where unable to be gathered due to the nature of failure for those directions. Stage 1 and 2 are linear regressions while stage 3 is a degree 2, to best match the curve. Linear regression is used on stage 2 because ultimately that represents the energy absorbed. With more data, a degree n x 2 regression would be a more appropriate regression for stage 2, where n is the number of layers. The table below is the approximation of each scenario and stage. Each stage starts and ends at the intersection of the next stage. There where two main objectives for these tests: firstly to understand how honeycomb performs under extreme compression, and secondly to be able to numerically represent the honeycomb structure. Both of these where completed to varying degrees. It is important understood how the honeycomb fails. If it is crushed in the T direction it retains its integrity even after being crushed, however when crushed in the L and W crush direction, if it fails, it disintegrates, and loses all integrity. Also, the more layers the more time the material spends in stage 2. Failure is started by buckling, thus if there is any imperfection, the stress will not spike but transition straight into stage 2. The data collected and aggregated can be used for initial simulation of honeycomb used in packages. The initial testing has set the ground work for more data to be collected in order to verify results and to allow more confidence in the simulation results. This is only the initial data collected. The next steps is to continue to collect more data to verify results and to test more variants of honeycomb, with different brazing, layers, and shape.

36 MATERIALS SCIENCE↗

Strategies for on-chip digital data compression for X-ray pixel detectors

Here, the continued desire for X-ray pixel detectors with higher frame rates will stress the ability of application-specific integrated circuit (ASIC) designers to provide sufficient off-chip bandwidth to reach continuous frame rates in the 1 MHz regime. To move from the current 10 kHz to the 1 MHz frame rate regime, ASIC designers will continue to pack as many power-hungry high-speed transceivers at the periphery of the ASIC as possible. In this paper, however, we present new strategies to make the most efficient use of the off-chip bandwidth by utilizing data compression schemes for X-ray photon-counting and charge-integrating pixel detectors. In particular, we describe a novel in-pixel compression scheme that converts from analog to digital converter units to encoded photon counts near the photon Poisson noise level and achieves a compression ratio of >1.5x independent of the dataset. In addition, we describe a simple yet efficient zero-suppression compression scheme called "zeromask" (ZM) located at the ASIC's edge before streaming data off the ASIC chip. ZM achieves average compression ratios of >4x, >7x, and >8x for high-energy X-ray diffraction, ptychography, and X-ray photon correlation spectroscopy datasets, respectively. We present the conceptual designs, register-transfer level block diagrams, and the physical ASIC implementation of these compression schemes in 65 nm CMOS. When combined, these two digital compression schemes could increase the effective off-chip bandwidth by a factor of 6-12x.

47 OTHER INSTRUMENTATION↗

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3↗