Precision Cascade: A novel algorithm for multi-precision extreme compression
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.
Handling large-scale scientific data in high-performance computing (HPC) environments poses significant challenges, including excessive I/O, high storage costs, and slow query performance. Traditional approaches often require full data decompression and scans, making them impractical for real-time or interactive analysis. To address these limitations, we introduce Eureka, a unified data-index co-compression framework that enables fine-grained access and efficient range queries on compressed scientific datasets. Eureka integrates spatial domain decomposition with block-wise error-bounded lossy compression to support selective decompression. It constructs a hierarchical AVL-tree index during compression to capture block-level value ranges, enabling fast pruning during query execution. To reduce metadata overhead, the index itself is also compressed while ensuring recall-preserving results. Experiments on six diverse HPC simulation datasets show that Eureka achieves up to 25x data compression and over 300x index compression, surpassing state-of-the-art compressors such as SZ3 and ZFP in rate-distortion performance. Additionally, Eureka delivers over 30x speedup for low-selectivity range queries, making it a scalable and efficient solution for modern scientific data analysis.
Methanol is a potentially attractive fuel for marine and off−road engines owing to its availability at bunkering and global distribution locations. Although methanol is well−distributed worldwide, its fuel chemistry and ignition properties make it poorly suited as a direct drop−in replacement for diesel fuel in compression−ignition engines. However, industrial processes are regularly used to convert methanol, via catalytic dehydration, to dimethyl ether (DME) over nonprecious metal catalysts. This chemical conversion can occur at relatively low pressures, temperatures, and catalyst space velocities, highlighting a potential opportunity to generate DME via onboard catalytic dehydration of methanol. DME’s fuel kinetic and ignition properties for compression ignition are much more favorable than those of methanol or even diesel fuel, but DME is more challenging than diesel fuel or methanol to pump, store, and deliver through conventional diesel fueling injection hardware. Thus, a potential opportunity exists to use the ignition and kinetic properties of DME, with the transportation and delivery advantages of methanol, in a methanol−fueled mixing−controlled compression−ignition engine. The present work explores performance, combustion behavior, and emissions reduction opportunities for methanol mixing−controlled combustion, enabled by a HCCI of DME that represents a small fraction of the total fuel energy that can be generated onboard via catalytic dehydration of methanol.
Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.
The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, largely through increasing the number of stations in the array and the frequency span sampled. The two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and the constraints imposed by I/O, as I/O bandwidths drop significantly on cold storage. An example of this is the data deluge expected from the SKA Telescopes of more than 60 PB per day, all to be stored on the buffer filesystem. While compressing the data is an obvious solution, the impacts on the final data products are hard to predict. In this paper, we chose an error-controlled compressor – MGARD – and applied it to simulated SKA-Mid and real pathfinder visibility data, in noise-free and noise-dominated regimes. As the data have an implicit error level in the system temperature, using an error bound in compression provides a natural metric for compression. MGARD ensures the compression incurred errors adhere to the user-prescribed tolerance. To measure the degradation of images reconstructed using the lossy compressed data, we proposed a list of diagnostic measures, exploring the trade-off between these error bounds and the corresponding compression ratios, as well as the impact on science quality derived from the lossy compressed data products through a series of experiments. We studied the global and local impacts on the output images for continuum and spectral line examples. We found relative error bounds of as much as 10%, which provide compression ratios of about 20, have a limited impact on the continuum imaging as the increased noise is less than the image RMS, whereas a 1% error bound (compression ratio of 8) introduces an increase in noise of about an order of magnitude less than the image RMS. For extremely sensitive observations and for very precious data, we would recommend a 0.1% error bound with compression ratios of about 4. These have noise impacts two orders of magnitude less than the image RMS levels. At these levels, the limits are due to instabilities in the deconvolution methods. We compared the results to the alternative compression tool DYSCO, in both the impacts on the images and in the relative flexibility. MGARD provides better compression for similar error bounds and has a host of potentially powerful additional features.
Wire Arc Additive Manufacturing (WAAM) is an advanced manufacturing technology which utilizes welding systems to generate three dimensional geometries in a layer-by-layer fashion. Distortion or warping of a print substrate and WAAM components due to thermally induced residual stresses is an ongoing challenge limiting the widespread adoption of WAAM technologies for producing components. In this manuscript, a novel approach is described to address thermal distortion in deposited components by applying lateral compressions along the length of the deposited material. To demonstrate this method, a series of single-track walls were printed and compressed at evenly spaced intervals using a modified hydraulic cutter tool. The jaws of the tool were modified to compress material rather than to shear it. A mathematical model was developed to relate the curvature of the deposited material to the volume of compression required to eliminate this distortion. Validation of this model was performed using 3D scan data to compare the change in wall curvature induced by compression to the volume of the applied compressions. Substrate deflection was also compared against a control wall, and implementation of wall compression reduced maximum deflections by 93% across a series of four depositions and subsequent compressions. Wall cross sections were also analyzed to determine the impact of compression on material hardness and grain structure. The results demonstrate that successively placed lateral compressions can effectively control and potentially eliminate bending distortion in printed parts. This methodology can be further developed to form a robust model for correction of thermally-induced distortion in WAAM components.
Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.
The construction of highly coherent X-ray sources, combined with next-generation detectors that are larger and faster, has enabled new research opportunities across the scientific landscape. Among the techniques that benefit most from these advancements is X-ray photon correlation spectroscopy (XPCS), where faster acquisition unlocks the ability to study faster dynamics within samples. However, faster acquisition on larger detectors also introduces unprecedented challenges for online data processing and offline data storage. Such challenges are particularly prominent for XPCS, where real time analyses require simultaneous calculation of all the previously acquired data in the time series. We present a homomorphic compression scheme to effectively reduce the computational time and memory space required for XPCS analysis. Leveraging similarities in the mathematical expression between a matrix-based compression algorithm and the correlation calculation, our approach allows direct operation on the compressed data without their decompression. The offline compression scheme extends storage capacity by a factor of 40 while preserving key features in the lossy compressed data. Meanwhile, the online compression scheme reduces the computational time to below 1 ms, enabling real time calculation of the correlation functions at kHz framerate. Our demonstration of a homomorphic compression of scientific data provides an effective solution to the big data challenge at coherent light sources. Beyond the example shown in this work, the framework can be extended to facilitate real-time operations directly on a compressed data stream for other techniques.
Carbon fiber provides opportunity to reduce weight in structural composites, including wind turbine blades, due to the material's superior specific stiffness and specific strength compared to alternatives. Despite these advantages, cost and compressive performance are considered weaknesses for carbon fiber products available today. Studies to produce low-cost carbon fiber alternatives, including the use of textile-derived precursor systems, have shown progress and merit through the DOE/ORNL low-cost carbon fiber initiatives. Here, this work focuses on enabling increases in compressive strength through design of the carbon fiber geometry, applicable to both textile and conventional precursor systems, while also providing opportunities to reduce carbon fiber processing costs. Fiber-resin interface and fiber alignment are among the most frequently cited factors controlling composite compressive performance. However, it is believed that there is opportunity in traditionally unexplored routes to increasing compressive strength through alteration of the carbon fiber geometry by increasing the fiber area moment of inertia and/or the fiber perimeter and interfacial area. This paper presents initial results from manufacturing carbon fiber materials to assess the impacts of carbon fiber size on tested composite compressive performance with projected neutral or even beneficial impact on fiber and composite manufacturing economics. Carbon fiber systems with increasing size illustrate a favorable correlation for compressive performance greater than predicted from a micromechanical failure model. The manufacturing and mechanical test results support the hypothesis of this work that alterations to fiber geometry can be used to produce improvements of the compressive strength of carbon fiber reinforced polymers and provide incentive for related work in designing alternative shapes to further enhance compressive performance.
In situ X-ray nanotomography experiments where tensile or compressive force is applied on the sample require specialized equipment. A compression-tension device with fluid flow-through capability has been designed for X-ray nanotomography beamlines. The compression-tension cell is equipped with a triaxial stage for sample alignment and a high sensitivity loadcell for measurement of applied force. To handle the <100 µm samples used for X-ray nanotomography imaging and for loading samples on the compression-tension cell a sample manipulator has been built. The sample manipulator is capable of selecting a single <100 µm particle for nanotomography scanning while viewing multiple samples under an optical microscope. To test the functionality of these two devices an initial compression experiment involving two glass beads was performed. To demonstrate instrument stability two spherical glass beads were compressed from a no load condition until one of the beads fractured. Nanotomography data were collected at each step of increasing compressive force. The experimentally observed contact area of the spherical glass beads was compared with the theoretical estimate using the Hertz analysis. To demonstrate the fluid flow capability, two calcite grains were compressed against each other under a calcite saturated solution. Surface topological changes were observed for the stressed grain contact area.
Many polyatomic astrophysical plasmas are compressible and out of chemical and thermal equilibrium, introducing a bulk viscosity into the plasma via the internal degrees of freedom of the molecular composition, directly impacting the decay of compressible modes, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$. This is especially important for small-scale, turbulent dynamo processes in the interstellar medium (ISM), which are known to be sensitive to the effects of compression. To control the viscous properties of $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$, we perform trans-sonic, visco-resistive dynamo simulations with additional bulk viscosity $\nu _{\text{bulk}}$, deriving a new $\nu _{\text{bulk}}$ Reynolds number $\text{Re}_{\text{bulk}}$, and viscous Prandtl number $\text{P}\nu \equiv \text{Re}_{\text{bulk}}/ \text{Re}_{\text{shear}}$, where $\text{Re}_{\text{shear}}$ is the shear viscosity Reynolds number. We derive a framework for decomposing $E_{\rm mag}$ growth rates into incompressible and compressible terms via orthogonal tensor decompositions of $\boldsymbol {\nabla }\otimes \mathrm{{\boldsymbol {\mathit {v}}}}$, where $\mathrm{{\boldsymbol {\mathit {v}}}}$ is the fluid velocity. We find that $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ play a dual role, growing and decaying $E_{\rm mag}$, and that field-line stretching is the main driver of growth, even in compressible dynamos. In the absence of $\nu _{\text{bulk}}$ ($\text{P}\nu \rightarrow \infty$), $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ pile up on small-scales, creating a spectral bottleneck, which disappears for $\text{P}\nu \approx 1$. As $\text{P}\nu$ decreases, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ are dissipated at increasingly larger scales, in turn suppressing incompressible modes through a coupling between high-k modes. We emphasize the importance of further understanding the role of $\nu _{\text{bulk}}$ in compressible astrophysical plasmas, which we estimate could be as strong as the shear viscosity in the cold ISM, and highlight that compressible direct numerical simulations without bulk viscosity have unresolved compressible mode dissipation scales.
Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.
Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).
To gain insights into thermodynamic states attained during shock compression of cemented tungsten carbide with 3.7 wt.% cobalt binder, we present results of longitudinal sound (release wave) speed measurements and their analysis at peak stresses up to 100 GPa (volumetric compression ratio ~ 15%). The sound speeds are determined using front-surface impact and release-wave overtake plate impact experimental configurations using laser interferometry. The measured sound speed data along with estimates for bulk sound speeds obtained using the fourth-order Birch-Murnaghan EoS and thermodynamics are used to determine the longitudinal moduli and shear moduli of shocked tungsten carbide at the various peak compression states attained in the experiments. Here, the longitudinal sound speeds were found to increase linearly with volume compression ratio from 6.97 ± 0.010 km/s at ambient conditions to 8.26 ± 0.156 km/s at a volume compression ratio of ~ 15%. The corresponding longitudinal elastic moduli also increase nearly linearly with the volume compression ratio but remain consistently lower than their theoretical predictions based on continuum models with no damage. Also, the sensitivity of shear moduli to pressure, as predicted by the Steinberg-Guinan model, is reduced substantially and the shear moduli of cemented WC with 3.7 wt.% Co remains nearly constant at ~ 310 GPa at the various peak compression stress states investigated in the present study.