Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compression techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

cuZ-Checker: A GPU-Based Ultra-Fast Assessment System for Lossy Compressions

Lossy compression is becoming an indispensable technique for the success of today's extreme-scale high-performance computing projects that produce vast volumes of data during scientific simulations or instrument data acquisitions. Comprehensively understanding the compression quality and performance of different lossy compressors is critical to selecting the best-fit compressors and using them properly and efficiently in practice. A few lossy compression assessment tools (e.g., Z-checker) have been developed, but none of them support the execution in a GPU environment. This is a significant gap because many recent extreme-scale applications and lossy compressors (e.g., cuSZ) can run entirely within GPUs. In this work, we develop an efficient lossy compression measuring system (called cuZ-Checker) on the GPU platform, which aims to perform the lossy compression quality and performance assessment completely within the GPU environment. Our contribution is threefold. (1) We develop a novel GPU-based lossy compression measuring framework using a computation pattern-based design approach. This approach classifies the computing-intensive metrics into three categories based on their patterns which creates large opportunities for kernel fusion and data reuse. (2) For each pattern in cuZ-Checker, we develop a CUDA kernel and provide fine-grained optimizations to boost its performance. (3) We thoroughly evaluate our cuZ-checker on a V100 GPU using four real-world scientific application datasets. Experiments show that cuZ-Checker can significantly accelerate the overall lossy compression assessment performance by 23X similar to 31X compared with the OpenMP-based multithreading CPU performance. To the best of our knowledge, this is the first lossy compression measuring system designed for GPU devices.

GPU↗

Micropillar Compression of Additively Manufactured 316L Stainless Steels after 2 MeV Proton Irradiation: A Comparison Study between Planar and Cross-Sectional Micropillars

A micropillar compression study with two different techniques was performed on proton-irradiated additively manufactured (AM) 316L stainless steels. The sample was irradiated at 360 °C using 2 MeV protons to 1.8 average displacement per atom (dpa) in the near-surface region. A comparison study with mechanical test and microstructure characterization was made between planar and cross-sectional pillars prepared from the irradiated surface. While a 2 MeV proton irradiation creates a relatively flat damage zone up to 12 µm, the dpa gradient by a factor of 2 leads to significant dpa uncertainty along the pillar height direction for the conventional planar technique. Cross-sectional pillars can significantly reduce such dpa uncertainty. From one single sample, three cross-sectional pillars were able to show dpa-dependent hardening. Furthermore, post-compression transmission electron microscopy allows the determination of the deformation mechanism of individual micropillars. Cross-sectional micropillar compression can be used to study radiation-induced mechanical property changes with better resolution and less data fluctuation.

transmission electron microscopy↗

Macro-level mechanical interlocking: A rapid joining approach for additively manufactured compression molded composite panels

Composite joining typically involves multiple steps, such as drilling and surface treatment, as part of the manufacturing process, which leads to low throughput and long cycle times. In the present study, we demonstrated a macro-level mechanical interlocking (MI) based, rapid joining technique to assemble additively manufactured compression molded (AMCM) panels, enabling the production of parts larger than the mold dimensions. Composite panels made of 20 wt% short carbon fiber reinforced acrylonitrile butadiene styrene (CF/ABS) were joined using MI features of various geometries, namely tree (TR), dovetail (Dov), rectangle 2 (Rect2), and rectangle 1 (Rect1), and their in-plane strength was evaluated. The resultant strength of the tested MI joints reached up to 74 % of the baseline tensile strength (i.e., the ‘no joint’ case). Observations from optical and scanning electron microscopy revealed inadequate polymer diffusion between the adherends, indicating that the joint strength was primarily derived from mechanical interlocking. Additionally, the fracture surfaces exhibited stress-whitening marks, which were characterized using differential scanning calorimetry (DSC). The increase in melting enthalpy suggested local stretching of polymer chains due to MI. Finite element analysis (FEA) indicated that the Rect1 MI feature, which generated the lowest stress concentration, outperformed the others in terms of joint strength, achieving 42 MPa. As a demonstration of the MI joining method, a battery box tray measuring 108 cm × 34 cm using a mold with an effective dimension of 36 cm × 34 cm successfully manufactured, resulting in a part with an area three times larger than the mold. In conclusion, this study presents a promising approach to improving composite joining techniques while minimizing production complexities.

In-plane joining↗

Exploring Autoencoder-based Error-bounded Compression for Scientific Data

Error-bounded lossy compression is becoming an indispensable technique for the success of today's scientific projects with vast volumes of data produced during the simulations or instrument data acquisitions. Not only can it significantly reduce data size, but it also can control the compression errors based on user-specified error bounds. Autoencoder (AE) models have been widely used in image compression, but few AE-based compression approaches support error-bounding features, which are highly required by scientific applications. To address this issue, we explore using convolutional autoencoders to improve error-bounded lossy compression for scientific data, with the following three key contributions. (1) We provide an in-depth investigation of the characteristics of various autoencoder models and develop an error-bounded autoencoder-based framework in terms of the SZ model. (2) We optimize the compression quality for main stages in our designed AE-based error-bounded compression framework, fine-tuning the block sizes and latent sizes and also optimizing the compression efficiency of latent vectors. (3) We evaluate our proposed solution using five real-world scientific datasets and comparing them with six other related works. Experiments show that our solution exhibits a very competitive compression quality from among all the compressors in our tests. In absolute terms, it can obtain a much better compression quality (100%similar to 800% improvement in compression ratio with the same data distortion) compared with SZ2.1 and ZFP in cases with a high compression ratio.

Liu, Jinyang↗

Revisiting Huffman Coding: Toward Extreme Performance on Modern GPU Architectures

Today's high-performance computing (HPC) applications are producing vast volumes of data, which are challenging to store and transfer efficiently during the execution, such that data compression is becoming a critical technique to mitigate the storage burden and data movement cost. Huffman coding is arguably the most efficient Entropy coding algorithm in information theory, such that it could be found as a fundamental step in many modern compression algorithms such as DEFLATE. On the other hand, today's HPC applications are more and more relying on the accelerators such as GPU on supercomputers, while Huffman encoding suffers from low throughput on GPUs, resulting in a significant bottleneck in the entire data processing. In this paper, we propose and implement an efficient Huffman encoding approach based on modern GPU architectures, which addresses two key challenges: (1) how to parallelize the entire Huffman encoding algorithm, including codebook construction, and (2) how to fully utilize the high memory-bandwidth feature of modern GPU architectures. The detailed contribution is fourfold. (1) We develop an efficient parallel codebook construction on GPUs that scales effectively with the number of input symbols. (2) We propose a novel reduction based encoding scheme that can efficiently merge the codewords on GPUs. (3) We optimize the overall GPU performance by leveraging the state-of-the-art CUDA APIs such as Cooperative Groups. (4) We evaluate our Huffman encoder thoroughly using six real-world application datasets on two advanced GPUs and compare with our implemented multi-threaded Huffman encoder. Experiments show that our solution can improve the encoding throughput by up to 5.0x and 6.8x on NVIDIA RTX 5000 and V100, respectively, over the state-of-the-art GPU Huffman encoder, and by up to 3.3x over the multi-thread encoder on two 28-core Xeon Platinum 8280 CPUs.

Tian, Jiannan↗

Revisiting Huffman Coding: Toward Extreme Performance on Modern GPU Architectures

Today's high-performance computing (HPC) applications are producing vast volumes of data, which are challenging to store and transfer efficiently during the execution, such that data compression is becoming a critical technique to mitigate the storage burden and data movement cost. Huffman coding is arguably the most efficient Entropy coding algorithm in information theory, such that it could be found as a fundamental step in many modern compression algorithms such as DEFLATE. On the other hand, today's HPC applications are more and more relying on the accelerators such as GPU on supercomputers, while Huffman encoding suffers from low throughput on GPUs, resulting in a significant bottleneck in the entire data processing. In this paper, we propose and implement an efficient Huffman encoding approach based on modern GPU architectures, which addresses two key challenges: (1) how to parallelize the entire Huffman encoding algorithm, including codebook construction, and (2) how to fully utilize the high memory-bandwidth feature of modern GPU architectures. The detailed contribution is fourfold. (1) We develop an efficient parallel codebook construction on GPUs that scales effectively with the number of input symbols. (2) We propose a novel reduction based encoding scheme that can efficiently merge the codewords on GPUs. (3) We optimize the overall GPU performance by leveraging the state-of-the-art CUDA APIs such as Cooperative Groups. (4) We evaluate our Huffman encoder thoroughly using six real-world application datasets on two advanced GPUs and compare with our implemented multithreaded Huffman encoder. Experiments show that our solution can improve the encoding throughput by up to 5.0× and 6.8× on NVIDIA RTX 5000 and V100, respectively, over the state-of-the-art GPU Huffman encoder, and by up to 3.3× over the multithread encoder on two 28-core Xeon Platinum 8280 CPUs.

Tian, Jiannan↗

Accelerating Lossy and Lossless Compression on Emerging BlueField DPU Architectures

Data compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios.

Li, Yuke↗

High-performance molded composites using additively manufactured preforms with controlled fiber and pore morphology

Here, large-scale multimaterial preforms produced by additive manufacturing (AM) underwent compression molding (CM) to produce high-performance thermoplastic composites reinforced with short carbon fibers. AM and CM techniques were integrated to control the fiber orientation (microstructure) and to reduce void content for the improved mechanical performance of the composite. The new integrated manufacturing technique is termed “additive manufacturing-compression molding” (AM-CM). For the present study, the most common materials were used for large-scale printing, i.e., acrylonitrile butadiene styrene (ABS), carbon fiber (CF)–filled ABS (CF/ABS) and glass fiber (GF)–filled ABS (GF/ABS). Three different manufacturing processes; (a) AM (b) extrusion compression molding (ECM), and (c) AM-CM were used to prepare four different panel configurations: (1) neat ABS, (2) CF/ABS, (3) overmold (CF/ABS over neat ABS), and (4) sandwich (neat ABS between two CF/ABS layers). The mechanical properties (tensile and flexural strength and modulus, and Izod impact energy) of samples prepared via all three manufacturing processes were compared. X-ray microcomputer tomography was employed to evaluate the fiber orientation distribution and the volumetric porosity content. The preform maintained high fiber alignment (≈ 82% of fibers within the range of 0–20° in the deposition direction), and the volumetric porosity was reduced by 50% from 3.79% to 1.91% after compression. The alignment of long pores along the deposition direction was also observed. The mechanical properties are discussed with correlation to the fiber alignment and void content in the samples. CF/ABS samples prepared by AM-CM showed significant improvement of 11.15%, 35.27%, 28.6%, and 74.3% in the tensile strength, tensile modulus, flexural strength, and flexural modulus, respectively, when compared with samples prepared by ECM. Unique aspects of this study are the demonstration of large-scale multimaterial AM and the use of multimaterials as preforms to make high-performance composites.

36 MATERIALS SCIENCE↗

OptZConfig: Efficient Parallel Optimization of Lossy Compression Configuration

Lossless compressors have very low compression ratios that do not meet the needs of today's large-scale scientific applications that produce vast volumes of data. Error-bounded lossy compression (EBLC) is considered a critical technique for the success of scientific research. Although EBLC allows users to set an error bound for the compression, users have been unable to specify the requirements on the compression quality, limiting practical use. Our contributions are: (1) We formulate the problem of configuring EBLC to preserve a user-defined metric as an optimization problem. This allows many classes of new metrics to be preserved, which improves over current practices. (2) We present a framework, OptZConfig, that can adapt to improvements in the search algorithm, compressor, and metrics with minimal changes, enabling future advancements in this area. (3) We demonstrate the advantages of our approach against the leading methods to configure compressors to preserve specific metrics. Here, our approach improves compression ratios against a specialized compressor by up to 3 x, has a 56x speedup over FRaZ, 1000x speedup over MGARD-QOI post tuning, and 110x speedup over systematic approaches which had not been bounded by compressors before.

97 MATHEMATICS AND COMPUTING↗

Full-Field Strain Measurement Integrated with Two Dimension Regression Analysis to Evaluate the Bi-Modulus Elastic Properties of Isotropic and Transversely Isotropic Materials

Background: Measuring the physical properties of shale is critical for optimizing engineering activities such as geothermal energy generation and hydraulic fracturing. Shale is a transversely isotropic material. Furthermore, this material can also include micro and macro cracks at different locations and orientations that cause it to behave differently under tensile or compressive loading. Objective: In this work, a combined experimental–numerical approach is proposed to evaluate the bi-modulus elastic properties of isotropic and transversely isotropic materials. Methods: Full-field strain measurements for a circular disk under diametral compression are integrated with a regression analysis technique to evaluate the elastic properties of bi-modulus materials subjected to tensile and compressive loads using two loading configurations on the same specimen. Digital Image Correlation (DIC) is used to measure the full-field strains. Subsequently, in the case of an isotropic material, a linear least-squares approach is utilized to process the experimentally determined strains in conjunction with analytical expressions of the stress fields (in terms of far-field loading) to determine the elastic modulus E, the shear modulus G, and the Poisson’s ratio $v$. In the case of a transversely isotopic material, such as shale, a finite element model is implemented to determine the stress fields (again in terms of far-field loading), which is followed by repeating the previous regression analysis in an iterative process to estimate the elastic parameters. Results: The results show that the proposed technique successfully provides a complete set of elastic properties as a function of both the loading condition and the principal material directions. The technique is validated by measurements on a known isotropic material and then applied to determine the properties of shale. Conclusion: In this work, the proposed approach is successfully used to calculate the bi-modulus elastic response of poly(methyl meth- acrylate) (PMMA) and shale. As expected, PMMA exhibits an isotropic response with no bi-modulus effect, however, shale exhibits both transverse isotropy and a bi-modulus effect. Therefore, this approach holds promise for investigating the elastic properties of materials like rocks and fiber-reinforced composite laminates as functions of the principal material directions and the loading conditions.

42 ENGINEERING↗

Dynamic signal recovery in distribution grids using compressive lossy measurements

Distribution system state estimation requires reliable aggregation of the measured data. However, the large volume of the measured data imposes a significant stress on the underlying communication infrastructure. With the challenges associated with measurement availability, current distribution systems are typically unobservable. To cope with the unobservability issue, compressive sensing theory allows us to recover system state information from a small number of measurements provided the states of the distribution system exhibit sparsity. In this paper, we evaluate the robustness of an updated Kalman filtered modified compressive sensing (KF-ModCS) technique that dynamically estimates the grid states using a small fraction of measured data. In practice, measurements used for sparsity based state estimation may also be intermittent due to communication network induced losses. Further, to understand the effect of packet losses on KF-ModCS, we provide an upper bound for the expected variances of the state estimation error for a given rate of information loss. This upper bound is further improved if the support set of the sparse signal that characterizes the state dynamics does not change over time and/or the reduced model is observable. Simulations based on two practical data sets collected from actual customers in a distribution grid validate the theoretical results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Design optimization of lightweight automotive seatback through additive manufacturing compression overmolding of metal polymer composites

With the growing demand for enhanced automotive fuel efficiency and environmental sustainability, there is a need for lightweighting automotive components through innovative design and manufacturing processes. Here, this study leverages a combination of numerical iterative design optimization and hybrid additive manufacturing–compression molding (AM-CM) technique for metal polymer composites to lightweight an automotive seatback. The AM-CM process enables robust mechanical interlocking between metals and composites, boasting high stiffness and strength with low overall density. Replacing metallic components with such metal polymer composites allows for comparable mechanical performance while significantly reducing the overall weight. First, the automotive seatback design space is reduced to critical load carrying regions using topology optimization and high stress concentration areas are identified using finite element analysis. Next, a lightweight metal polymer subcomponent is designed for a high stress concentration region. The full seatback frame with spatially heterogeneous material-specific design is then iteratively optimized to enable enhanced stiffness with minimal weight. Overall, the automotive seatback frame designed with location-specific metal, polymer, and metal polymer composite materials weighs 20% less than the metal-only design while exhibiting similar stiffness.

36 MATERIALS SCIENCE↗

Exploring scenarios for enhanced fuel compression and performance on the National Ignition Facility with machine-learning-aided design techniques

Recent fusion experiments on the National Ignition Facility (NIF) have achieved ignition, producing multi-MJ fusion yields for input laser energies of roughly 2 MJ [Abu-Shawareb et al., Phys. Rev. Lett. 132, 065102 (2024)]. Building on the success of the target designs that have achieved ignition, we explore new implosion scenarios predicted to generate significantly more compression of the dense DT ice layer and correspondingly higher yields while preserving many of the key physics characteristics of present-day ignition designs. Our main result is a novel 3-shock implosion scheme that effectively minimizes the shock-induced entropy in the dense, accelerating DT shell and maximizes the resulting fuel compression subject to a fixed leading shock strength consistent with present-day ignition experiments, which is necessary to melt the crystalline high-density carbon ablator. Compared to the first NIF experiment to fulfill Lawson's ignition criterion, shot N210808 [Abu-Shawareb et al., Phys. Rev. Lett. 129, 075001 (2022)], our design exhibits a 40% increase in simulated peak areal density (ρR) and a 5× increase in 1D fusion yield using a 4% lighter ablator and identical DT payloads. We also present a complete integrated 2D hohlraum design and laser pulse specifications capable of generating the desired 3-shock drive and maintaining control of the low-mode capsule implosion symmetry, where the increase in simulated 2D yield relative to N210808 is > 10×. This new implosion regime was discovered with help from a machine-learning-enabled capsule design optimization framework. We outline the workflow this automated tool uses to identify improved design candidates by running several rounds of capsule simulations, constructing a surrogate model mapping input variations to key physics output quantities, and querying the resulting statistical model to propose adjustments to the x-ray drive and capsule to reach a set of physics objectives prescribed by the designer.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Spike-and-Slab Shrinkage Priors for Structurally Sparse Bayesian Neural Networks

Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.

97 MATHEMATICS AND COMPUTING↗

Effect of recycled fibers and shredded intermediates variation on the mechanical properties and energy absorption of fiber‐reinforced composite panels

Abstract The global composite industry generates large quantities of waste which mostly ends up in landfills due to a lack of established end‐use applications for multiple waste streams. The scrap from end‐of‐life (EoL) includes manufacturing waste such as dry chopped fiber tows, loose fibers, shredded fibers from fabric textile operations, cured/semi‐cured prepregs, and fully cured composite structure waste from aircraft, automobiles, wind blades, boats, and pressure vessels. In this work, different composite waste streams were reduced to shredded intermediates, followed by simple blending, and subjected to wet compression molding to produce composite panels. The panels/plaques were tested for mechanical properties (flexure and impact), fiber‐matrix wet‐out, and property bounds. It was found that wet‐compression molding was a viable and scale‐able approach to produce recycled panels from EoL composites shredded scrap. Furthermore, full‐scale size panels for use in truck bodies and intermodal shipping container flooring were manufactured and their impact resistance was tested using a drop weight impact test. They were tested both for high‐ and low‐velocity load. In the case of high‐velocity load, the average impact load was 14,673 N; the average absorbed energy was 101.6 J; the average elastic energy was 11.7 J and the impact resistance was 1065 J/m. In the case of the low‐velocity drop weight impact test, it was found that the average impact load was 7877.358 N; the average absorbed energy was 9.718 J; the average elastic energy was 9.14 J, and the impact resistance was 184.1 J/m. The shredded composite was shown to be a candidate material for the manufacture of truck bodies and intermodal containers’ flooring panels. Highlights By using shredded intermediates from different composite waste streams, it is possible to manufacture composite panels. Wet–compression process is an appropriate technique for manufacturing recycled fiber composite panels. Regardless of the source of scrap, the mechanical properties of the produced composite panels were improved. The recycling process technology can be transformed to commercial scale to produce full‐size transportation flooring panels. A product pathway is established in consideration of lower cost and improved recyclability.

Vaidya, Uday↗

Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing

The Joint Laboratory on Extreme-Scale Computing (JLESC) was initiated at the same time lossy compression for scientific data became an important topic for the scientific communities. The teams involved in the JLESC played and are still playing an important role in developing the research, techniques, methods, and technologies making lossy compression for scientific data a key tool for scientists and engineers. Here, in this paper, we present the evolution of lossy compression for scientific data from 2015, describing the situation before the JLESC started, the evolution of this discipline in the past 8 years (until 2023) through the prism of the JLESC collaborations on this topic and some of the remaining open research questions.

Compression for AI↗

Boron‐polymer composites engineered for compression molding, foaming, and additive manufacturing

Abstract Boron (specifically 10 B) is the element of choice to shield thermal neutrons due to its large (n, α) cross‐section; however, very few polymer composites containing high boron concentrations are available. This study aimed to determine the maximum possible amount of boron that could be introduced into a polymer matrix. Diverse manufacturing techniques, ranging from additive manufacturing to compression molding, were employed to fabricate inks and filaments for 3D printing, foams, and flexible pads. Composites using siloxanes, poly(lactic acid), and acrylonitrile butadiene styrene containing up to 80 wt% boron were sucessufully fabricated. The addition of known plasticizers (polyethylene glycol) and reinforcing agents (carbon nanofibers and fumed silica) helped to overcome fabrication problems such as clogging of the printing nozzle or crumbling of compression molded parts. In addition, the thermal‐mechanical properties of these novel boron composites were determined and shown to vary according to boron concentration, presence of additives, and fabrication techniques utilized.

36 MATERIALS SCIENCE↗