Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING

ZFP: A compressed array representation for numerical computations

HPC trends favor algorithms and implementations that reduce data motion relative to FLOPS. We investigate the use of lossy compressed data arrays in place of traditional IEEE floating point arrays to store the primary data of calculations. Simulation is fundamentally an exercise in controlled approximation, and error introduced by finite-precision arithmetic (or lossy compression) is just one of several sources of error that need to be managed to ensure sufficient accuracy in a computed result. We describe ZFP, a compressed numerical format designed for in-memory storage of multidimensional arrays, and summarize theoretical results that demonstrate that the error of repeated lossy compression can be bounded and controlled. Furthermore, we establish a relationship between grid resolution and compression-induced errors and show that, contrary to conventional floating point, ZFP reduces finite-difference errors with finer grids. We present example calculations that demonstrate data reduction by 4x or more with negligible impact on solution accuracy. Our results further demonstrate several orders-of-magnitude increase in accuracy using ZFP over IEEE floating point and Posits for the same storage budget.

Lindstrom, Peter

Integrated photonic encoder for low power and high-speed image processing

Abstract Modern lens designs are capable of resolving greater than 10 gigapixels, while advances in camera frame-rate and hyperspectral imaging have made data acquisition rates of Terapixel/second a real possibility. The main bottlenecks preventing such high data-rate systems are power consumption and data storage. In this work, we show that analog photonic encoders could address this challenge, enabling high-speed image compression using orders-of-magnitude lower power than digital electronics. Our approach relies on a silicon-photonics front-end to compress raw image data, foregoing energy-intensive image conditioning and reducing data storage requirements. The compression scheme uses a passive disordered photonic structure to perform kernel-type random projections of the raw image data with minimal power consumption and low latency. A back-end neural network can then reconstruct the original images with structural similarity exceeding 90%. This scheme has the potential to process data streams exceeding Terapixel/second using less than 100 fJ/pixel, providing a path to ultra-high-resolution data and image acquisition systems.

47 OTHER INSTRUMENTATION

Lifting MGARD: Construction of (pre)wavelets on the interval using polynomial predictors of arbitrary order

MGARD (MultiGrid Adaptive Reduction of Data) is an algorithm for compressing and refactoring scientific data, based on the theory of multigrid methods. The core algorithm is built around stable multilevel decompositions of conforming piecewise linear $C^0$ finite element spaces, enabling accurate error control in various norms and derived quantities of interest. In this work, we extend this construction to arbitrary order Lagrange finite elements $\mathbb{Q}_p$, $p \geq 0$, and propose a reformulation of the algorithm as a lifting scheme with polynomial predictors of arbitrary order. Additionally, a new formulation using a compactly supported wavelet basis is discussed, and an explicit construction of the proposed wavelet transform for uniform dyadic grids is described.

Reshniak, Viktor [Oak Ridge National Laboratory (O

What to Support When You’re Compressing

Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.

Error-Bounded Lossy Compression

StOKeDMD: Streaming Occupation kernel dynamic mode decomposition

Dynamic mode decomposition (DMD) has become a common technique for constructing surrogate models for dynamical systems from observed system states. The Occupation Kernel DMD (OKDMD) method proposed in (Rosenfeld et al., 2022) and (Rosenfeld et al., 2024) is a Liouville operator based method that builds surrogate models from system state trajectories. Here, this paper proposes an extension of OKDMD to the case when the system states are observed in a streaming fashion, i.e., only a small fraction of the state trajectory is available at a given time. The developed method, Streaming Occupation Kernel DMD (StOKeDMD), accommodates the streaming data input by leveraging properties of specific choices of kernel functions and occupation kernels. We apply the StoKeDMD method as a compression method for streaming data, analyze the memory complexity, and demonstrate the performance of StoKeDMD in the compression of streaming data generated from a Lorenz system and a fluid flow simulation.

97 MATHEMATICS AND COMPUTING

Optimising the processing and storage of visibilities using lossy compression

The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, largely through increasing the number of stations in the array and the frequency span sampled. The two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and the constraints imposed by I/O, as I/O bandwidths drop significantly on cold storage. An example of this is the data deluge expected from the SKA Telescopes of more than 60 PB per day, all to be stored on the buffer filesystem. While compressing the data is an obvious solution, the impacts on the final data products are hard to predict. In this paper, we chose an error-controlled compressor – MGARD – and applied it to simulated SKA-Mid and real pathfinder visibility data, in noise-free and noise-dominated regimes. As the data have an implicit error level in the system temperature, using an error bound in compression provides a natural metric for compression. MGARD ensures the compression incurred errors adhere to the user-prescribed tolerance. To measure the degradation of images reconstructed using the lossy compressed data, we proposed a list of diagnostic measures, exploring the trade-off between these error bounds and the corresponding compression ratios, as well as the impact on science quality derived from the lossy compressed data products through a series of experiments. We studied the global and local impacts on the output images for continuum and spectral line examples. We found relative error bounds of as much as 10%, which provide compression ratios of about 20, have a limited impact on the continuum imaging as the increased noise is less than the image RMS, whereas a 1% error bound (compression ratio of 8) introduces an increase in noise of about an order of magnitude less than the image RMS. For extremely sensitive observations and for very precious data, we would recommend a 0.1% error bound with compression ratios of about 4. These have noise impacts two orders of magnitude less than the image RMS levels. At these levels, the limits are due to instabilities in the deconvolution methods. We compared the results to the alternative compression tool DYSCO, in both the impacts on the images and in the relative flexibility. MGARD provides better compression for similar error bounds and has a host of potentially powerful additional features.

Techniques: interferometric

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator

Sound speed and Grüneisen parameter up to three terapascal in shock-compressed iron

This paper presents the first sound speed and Grüneisen parameter data for fluid iron compressed to 3 TPa (30 million atmospheres) and 20 g/cm 3 on the Hugoniot. Both the sound speed and Grüneisen parameter are derivatives of the equation of state (EOS), and thus tightly constrain the contours of the EOS surface. The sound speed data are systematically lower than expected from a simple extrapolation of previous data. The Grüneisen parameter shows a 30% drop at pressures and temperatures above the melt transition. Furthermore, while some models compare well with either the sound speed or Grüneisen parameter, none of today’s state-of-the-art models can explain both sets of data. Furthermore these new data will provide pivotal benchmarks for both future theoretical EOSs of warm dense iron and modeling planetary states and processes.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Effect of Part Size, Displacement Rate, and Aging on Compressive Properties of Elastomeric Parts of Different Unit Cell Topologies Formed by Vat Photopolymerization Additive Manufacturing

Due to its ability to achieve geometric complexity at high resolution and low length scales, additive manufacturing (AM) has increasingly been used for fabricating cellular structures (e.g., foams and lattices) for a variety of applications. Specifically, elastomeric cellular structures offer tunability of compliance as well as energy absorption and dissipation characteristics. However, there are limited data available on compression properties for printed elastomeric cellular structures of different designs and testing parameters. In this work, the authors evaluate how unit cell topology, part size, the rate of compression, and aging affect the compressive response of polyurethane-based simple cubic, body-centered, and gyroid structures formed by vat photopolymerization AM. Finite element simulations incorporating hyperelastic and viscoelastic models were used to describe the data, and the simulated results compared well with the experimental data. Of the designs tested, only the parts with the body-centered unit cell exhibited differences in stress–strain responses at different part sizes. Of the compression rates tested, the highest displacement rate (1000 mm/min) often caused stiffer compressive behavior, indicating deviation from the quasi-static assumption and approaching the intermediate rate response. The cellular structures did not change in compression properties across five weeks of aging time, which is desirable for cushioning applications. This work advances knowledge on the structure–property relationships of printed elastomeric cellular materials, which will enable more predictable compressive properties that can be traced to specific unit cell designs.

36 MATERIALS SCIENCE

Modulated Thermomechanical Analysis of Compression-Molded High-Density Polyethylene

Thermomechanical analysis (TMA) experiments conducted on high-density polyethylene (HDPE) show both reversible and irreversible dimensional changes. To further explore these reversible and irreversible processes, modulated thermomechanical analysis (MTMA) was used. Before reliable data on compression-molded HDPE was collected, a parameter optimization was performed to obtain a suitable MTMA method. Once a suitable method was obtained, several MTMA experiments were conducted on compression-molded HDPE. This work highlights the steps taken during the MTMA parameter optimization and the results obtained from MTMA experiments conducted on pristine compression-molded HDPE samples.

36 MATERIALS SCIENCE

Compressive Response and Energy Absorption of Additively Manufactured Elastomers with Varied Simple Cubic Architectures

Additive manufacturing, and particularly the vat photopolymerization process, enables the fabrication of complex geometries at high resolution and small length scales, making it well-suited for fabricating cellular structures (e.g., foams and lattices). Among these, elastomeric cellular structures are of growing interest due to their tunable compliance and energy dissipation. However, comprehensive data on the compressive behavior of these structures remains limited, especially for investigating the structure-property effects from changing the density and distribution of material within the cellular structure. This study explores how the mechanical response of polyurethane-based simple cubic structures changes when varying volume fraction, unit cell length, and unit cell patterning, which have not been systematically investigated previously in additively manufactured elastomers. Increasing volume fraction from 10% to 50% yielded significant changes in compressive stress–strain performance (decreasing strain at 0.5 MPa by 41.6% and increasing energy absorption density by 3962.5%). Although changing the unit cell length between 2.5 and 7 mm in ~30 mm parts did not result in statistically different stress–strain responses, modifying the configuration of struts of different thicknesses across designs with 30% volume fraction altered the stress–strain behavior (differences of 12.5% in strain at 0.5 MPa and 109.4% for energy absorption density). Power law relationships were developed to understand the interactions between volume fraction, unit cell length, and elastic modulus, and experimental data showed strong fits (R 2 > 0.91). These findings enhance the understanding of how multiple structural design aspects influence the performance of elastomeric cellular materials, providing a foundation for informing strategic design of tailorable materials for diverse mechanical applications.

36 MATERIALS SCIENCE

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science

A surprising proliferation of detwinning in β -tin at extreme loading rates

Integrating data from dynamic compression experiments of condensed matter across three national laboratories has led to insight and quantitative calibration of materials strength over decades of loading rate. For many materials, a single strength model (such as PTW) is sufficient to capture the flow-stress strain rate relationship which is monotonic. Here, we show here that β -tin, a tetragonal metal, exhibits dramatic deviations from this behavior. Naive fitting to a single PTW model is insufficient to capture the behavior; indeed, such resulting inferred flow stress versus strain exhibits a non-monotonic behavior. We suggest a resolution to this by proposing that in β -tin there are important Bauschinger effects arising from favorable conditions for twinning and detwinning. A simple yield surface model when paired with PTW hardening captures the experimental data.

36 MATERIALS SCIENCE

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Advanced Polymer Characterization: Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)

Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry encodes structural information across diverse homo- and copolymer ensembles, yet decrypting these spectra requires a systematic analytical approach. We introduce Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)─a general cipher algorithm that applies modular arithmetic to filter monomer-derived mass contributions and cluster MALDI peaks by nonconstitutional repeating units (non-CRUs). MOSAIC performs sequential modular operations using monomer mass differences as base units to compress complex spectral data, revealing end-group distributions and comonomer incorporation. As a demonstration, we applied MOSAIC to five copolymers formed by two different polymerization mechanisms. Furthermore, the resulting remainder–mass plots clearly resolve polymer homologs with distinct non-CRUs into visually apparent clusters, enabling intuitive assignment of mass spectral features.

Wang, Hanlin M. [University of Illinois at Urbana−

Tuning the Interpolation Basis in a Multigrid Decomposition for Local Error Control

In the compression of scientific data, error-controlled compressors enable to considerably decrease the size of the dataset while maintaining adequate levels of accuracy. In this paper, we note that multi-level refactoring scheme such as MGARD i) rely on an approximation of the data based on the interpolation of coefficients, ii) estimate the resulting error with global metrics on the dataset. To improve on these two aspects, we propose a method that aims to divide the original dataset into blocks based on their smoothness and refactors each block separately with the most relevant interpolation order. We show the relevance of such a method on tailored datasets and the benefits and challenges when applying it to large scientific data.

Vidal, Nicolas [ORNL]