Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data compression techniques”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Real-time and post-hoc compression for data from Distributed Acoustic Sensing

Distributed Acoustic Sensing (DAS) is an emerging sensing technology that records the strain-rate along fiber optic cables at high spatial and temporal resolution. This technique is becoming a popular tool in seismology, hydrology, and other subsurface monitoring applications. However, due to the large coverage (10’s of km) and high density of measurements (1m spacing at 100’s of Hz), a DAS installation could produce terabytes of data records per day. Because many DAS instruments are deployed in remote locations, this large data size poses significant challenges to its transfer and storage. In this paper, we explore lossless compression methods to reduce the storage requirement in both real-time and post-hoc scenarios. Here we propose a two-stage compression method to improve the compression ratio and compression speed. This two-stage compression method could reduce the storage requirement by 40%, which is 20% more than other lossless methods, such as ZSTD. We demonstrate that the compression method could complete its operation well before the DAS instrument needs to output the next file, making it suitable for real-time DAS acquisition. We also implement a parallel compression method for a post-hoc scenario and demonstrate that our method could effectively utilize a parallel computer. With 256 CPU cores, our parallel compression method achieves the speed of 26GB/second.

58 GEOSCIENCES↗

A Tailored Convolutional Neural Network for Nonlinear Manifold Learning of Computational Physics Data Using Unstructured Spatial Discretizations

In this work, we propose a nonlinear manifold learning technique based on deep convolutional autoencoders that is appropriate for model order reduction of physical systems in complex geometries. Convolutional neural networks have proven to be highly advantageous for compressing data arising from systems demonstrating a slow-decaying Kolmogorov n-width. However, these networks are restricted to data on structured meshes. Unstructured meshes are often required for performing analyses of real systems with complex geometry. Our custom graph convolution operators based on the available differential operators for a given spatial discretization effectively extend the application space of deep convolutional autoencoders to systems with arbitrarily complex geometry that are typically discretized using unstructured meshes. We propose sets of convolution operators based on the spatial derivative operators for the underlying spatial discretization, making the method particularly well suited to data arising from the solution of partial differential equations. We demonstrate the method using examples from heat transfer and fluid mechanics and show better than an order of magnitude improvement in accuracy over linear methods.

97 MATHEMATICS AND COMPUTING↗

Super Resolving Unrolled Neural Networks for Remote Sensing

In remote sensing systems, the capabilities of the system are constrained by the complex interactions between size, weight, and power (SWAP) of potential designs. In electro-optical (EO) systems, examples of these critical parameters include the system’s sensitivity and resolution. Those parameters can be increased by ever larger optical apertures and focal planes but at the cost of more SWAP. Multi-image super resolution (MISR) techniques allow resolution to be enhanced via computation rather than more sophisticated optical hardware. These algorithms combine multiple images together into a single, higher resolution image, trading temporal resolution and computation for spatial resolution. Fielded MISR techniques, such as Drizzle, can require several hundred images to create a single super resolved image, implying reduced temporal resolution, increased data acquisition load, and limiting mission applications. Iterative techniques, such as model-based image reconstruction and compressive sensing, have been shown to create super resolved images using fewer images than Drizzle. They do this by posing an optimization problem that balances accuracy between a highly accurate physical model and an image model. In the case of super resolution, the physical model is defined by the relation between low resolution input images and the desired high resolution output image. The image model encodes some assumptions about the super resolved image. These assumptions are meant to suppress reconstruction artifacts that arise due to deterministic physical model error, stochastic measurement noise, and potential undersampling. In practice, the performance of iterative methods are limited by imaging models compatible with optimization. Deep learning-based methods can effectively learn image models of arbitrary complexity, but lack the theoretical explainability and robustness of iterative techniques. Consensus equilibrium (CE) generalizes the iterative techniques beyond optimization, enabling blackbox algorithms such as traditional and neural image denoisers to be used as the image model. CE-based approaches retain much of the explainability and robustness of iterative techniques while allowing the expressiveness of machine learning image models to be used. Additionally, by unrolling iterations of CE with an embedded image denoiser, the image denoiser can be further trained and specialized to the specific application with potentially higher quality reconstructions. Under this project, we demonstrated the feasibility of training an unrolled neural network based upon CE. While we didn’t train one, we showed that the CE process is differentiable and its gradient can be tractably computed. We also explored the usage of a variants of CE akin to generative neural works. Most importantly, we applied the CE framework to a number of problems including non-blind deconvolution, upsampling, single-image super resolution, MISR, event-based sensing, and saturated deconvolution. Our MISR prototype creates high quality reconstructions with an order of magnitude fewer images than previous approaches and, critically, produces these reconstructions fast enough for practical usage.

47 OTHER INSTRUMENTATION↗

Error-Bounded Learned Scientific Data Compression with Preservation of Derived Quantities

Scientific applications continue to grow and produce extremely large amounts of data, which require efficient compression algorithms for long-term storage. Compression errors in scientific applications can have a deleterious impact on downstream processing. Thus, it is crucial to preserve all the “known” Quantities of Interest (QoI) during compression. To address this issue, most existing approaches guarantee the reconstruction error of the original data or primary data (PD), but cannot directly control the problem of preserving the QoI. In this work, we propose a physics-informed compression technique that is composed of two parts: (i) reduction of the PD with bounded errors and (ii) preservation of the QoI. In the first step, we combine tensor decompositions, autoencoders, product quantizers, and error-bounded lossy compressors to bound the reconstruction error at high levels of compression. In the second step, we use constraint satisfaction post-processing followed by quantization to preserve the QoI. To illustrate the challenges of reducing the reconstruction errors of the PD and QoI, we focus on simulation data generated by a large-scale fusion code, XGC, which can produce tens of petabytes in a single day. The results show that our approach can achieve a high compression amount while accurately preserving the QoI within scientifically acceptable bounds.

97 MATHEMATICS AND COMPUTING↗

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

Physics-Driven Convolutional Autoencoder Approach for CFD Data Compressions: Preprint

With the growing size and complexity of turbulent flow models, data compression approaches are of the utmost importance to analyze, visualize, or restart the simulations. Recently, in-situ autoencoder-based compression approaches have been proposed and shown to be effective at producing reduced representations of turbulent flow data. However, these approaches focus solely on training the model using point-wise sample reconstruction losses that do not take advantage of the physical properties of turbulent flows. In this paper, we show that training autoencoders with additional physics-informed regularizations, e.g., enforcing incompressibility and preserving enstrophy, improves the compression model in three ways: (i) the compressed data better conform to known physics for homogeneous isotropic turbulence without negatively impacting point-wise reconstruction quality, (ii) inspection of the gradients of the trained model uncovers changes to the learned compression mapping that can facilitate the use of explainability techniques, and (iii) as a performance byproduct, training losses are shown to converge up to 12x faster than the baseline model.

auto-encoders↗

Simultaneous compression and opacity data from time-series radiography with a Lagrangian marker

Time-resolved radiography can be used to obtain absolute shock Hugoniot states by simultaneously measuring at least two mechanical parameters of the shock, and this technique is particularly suitable for one-dimensional converging shocks where a single experiment probes a range of pressures as the converging shock strengthens. However, at sufficiently high pressures, the shocked material becomes hot enough that the x-ray opacity falls significantly. Additionally, if the system includes a Lagrangian marker such that the mass within the marker is known, this additional information can be used to constrain the opacity as well as the Hugoniot state. In the limit that the opacity changes only on shock heating, and not significantly on subsequent isentropic compression, the opacity of the shocked material can be determined uniquely. More generally, it is necessary to assume the form of the variation of opacity with isentropic compression or to introduce multiple marker layers. Alternatively, assuming either the equation of state or the opacity, the presence of a marker layer in such experiments enables the non-assumed property to be deduced more accurately than from the radiographic density reconstruction alone. An example analysis is shown for measurements of a converging shock wave in polystyrene at the National Ignition Facility.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Efficient Clustering of Software Vulnerabilities using Self Organizing Map (SOM)

The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.

Panchal, Khyati↗

A Novel Deep Reinforcement Learning Approach to Traffic Signal Control with Connected Vehicles

The advent of connected vehicle (CV) technology offers new possibilities for a revolution in future transportation systems. With the availability of real-time traffic data from CVs, it is possible to more effectively optimize traffic signals to reduce congestion, increase fuel efficiency, and enhance road safety. The success of CV-based signal control depends on an accurate and computationally efficient model that accounts for the stochastic and nonlinear nature of the traffic flow. Without the necessity of prior knowledge of the traffic system’s model architecture, reinforcement learning (RL) is a promising tool to acquire the control policy through observing the transition of the traffic states. In this paper, we propose a novel data-driven traffic signal control method that leverages the latest in deep learning and reinforcement learning techniques. By incorporating a compressed representation of the traffic states, the proposed method overcomes the limitations of the existing methods in defining the action space to include more practical and flexible signal phases. The simulation results demonstrate the convergence and robust performance of the proposed method against several existing benchmark methods in terms of average vehicle speeds, queue length, wait time, and traffic density.

42 ENGINEERING↗

Concurrent measurement of strain and chemical reaction rates in a calcite grain pack undergoing pressure solution: Evidence for surface-reaction controlled dissolution

Pressure solution is inferred to be a significant contributor to sediment compaction and lithification, especially in carbonate sediments. For a sediment deforming primarily by pressure solution, the compaction rate should be directly related to the rate of calcite dissolution, transport along grain contacts, and calcite reprecipitation. Previous experimental work has shown that there is evidence that deformation in wet calcite grain packs is consistent with control by pressure solution, but considerable ambiguity remains regarding the rate limiting mechanism. We present the results of laboratory compaction experiments designed to directly measure calcite dissolution and precipitation rates (recrystallization rates) concurrently with strain rate to test whether measured rates are consistent with predicted rates both in absolute magnitude and time evolution. Recrystallization rates are measured using trace element chemistry (Sr/Ca, Mg/Ca) and isotopes (87Sr/86Sr) of fluids flowing slowly through a compacting grain pack as it is being triaxially compressed. Imaging techniques are used to characterize the grain contacts and strain effects in the post-experiment grain pack. Our data show that calcite recrystallization rates calculated from all three geochemical parameters are in approximate agreement and that the rates closely track strain rate. The geochemically inferred rates are close to predicted rates in absolute magnitude. Uncertainty in grain contact dimensions makes distinguishing between surface reaction control and diffusion control difficult. Measured reaction rates decrease faster than predicted from standard pressure solution creep flow laws. This inconsistency may indicate that calcite dissolution rates at grain contacts are more complex, and more time-dependent, than suggested by geometric models designed to predict grain contact stresses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of Compressibility Effects and Nonlinear Property Variations in a Supercritical CO2 Mixing Layer

Calculations are performed using the large eddy simulation technique to qualitatively assess the effects of fluid compressibility and thermodynamic nonlinearities on the dynamics of a supercritical flow field. Resulting data are also used to interrogate instantaneous subfilter-scale fields, as modeled by the mixed-dynamic Smagorinsky closure and gradient diffusion model. A three-dimensional wall-resolved calculation of a spatially evolving mixing layer composed solely of carbon dioxide is performed. This is followed by a coarse grid calculation to examine model performance as a function of resolution. Theoretical analysis indicates that the partial derivatives of density with respect to pressure and temperature are important modulators of the temperature and pressure field evolution, respectively. Results pertaining to subfilter velocity-velocity and velocity-temperature fields indicate that the models are most stressed in the regions where compressibility and thermo-fluid nonlinearities dominate the physics.

42 ENGINEERING↗

DeepAdversaries: examining the robustness of deep learning models for galaxy morphology classification

With increased adoption of supervised deep learning methods for work with cosmological survey data, the assessment of data perturbation effects (that can naturally occur in the data processing and analysis pipelines) and the development of methods that increase model robustness are increasingly important. In the context of morphological classification of galaxies, we study the effects of perturbations in imaging data. In particular, we examine the consequences of using neural networks when training on baseline data and testing on perturbed data. We consider perturbations associated with two primary sources: (a) increased observational noise as represented by higher levels of Poisson noise and (b) data processing noise incurred by steps such as image compression or telescope errors as represented by one-pixel adversarial attacks. We also test the efficacy of domain adaptation techniques in mitigating the perturbation-driven errors. We use classification accuracy, latent space visualizations, and latent space distance to assess model robustness in the face of these perturbations. For deep learning models without domain adaptation, we find that processing pixel-level errors easily flip the classification into an incorrect class and that higher observational noise makes the model trained on low-noise data unable to classify galaxy morphologies. On the other hand, we show that training with domain adaptation improves model robustness and mitigates the effects of these perturbations, improving the classification accuracy up to 23% on data with higher observational noise. Domain adaptation also increases up to a factor of ${\approx}2.3$ the latent space distance between the baseline and the incorrectly classified one-pixel perturbed image, making the model more robust to inadvertent perturbations. Successful development and implementation of methods that increase model robustness in astronomical survey pipelines will help pave the way for many more uses of deep learning for astronomy.

79 ASTRONOMY AND ASTROPHYSICS↗

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression↗

Initial Testing of an In Situ Load Retention Aging Vessel

A thermal aging vessel instrumented with load cells was fabricated. The primary function of the vessel is to continuously monitor the in situ load retention of up to three compressed polymer coupons undergoing thermally accelerated aging under nitrogen. A secondary function is to enable gas sampling of the vessel headspace during thermal aging. Heating of the vessel is achieved using a custom heater jacket. To improve upon our conventional aging study methods which require periodic interruption of aging to perform load testing in an Instron machine at room temperature, this technology aims to automate/facilitate data acquisition/analysis, improve data quality, and enable uninterrupted compression of the polymer which represents the service condition. As an example case to assess functionality of the in situ vessel, the load retention of a siloxane elastomer material additively manufactured by direct-ink-writing (DIW) was measured at three different isothermal aging temperatures for ~1 month. Initial compression of the coupons while near the aging temperature was achieved by temporarily opening the heated vessel to access the interior chamber and manually tightening four nuts to drive the heated compression plate down onto the heated coupons. Initial testing demonstrated achievement of the primary load retention monitoring function. Unfortunately, the vessel leaked which prevented gas sampling; an active purge was used to maintain a nitrogen atmosphere. Welded or otherwise sealed joints, which could be implemented in a future design, would likely eliminate leak paths. To apply time-temperature superposition (TTS), a technique used to provide long-term prediction of the load retention from short-term isothermal data, the load retention needed to be calculated relative to the load at an estimated “equilibrium” time, after most of the transient viscoelastic physical relaxation occurred. The peak load immediately after compression could not be used as the load retention basis for two reasons: (1) age-related changes must be isolated from non-age-related physical relaxation before applying TTS and (2) the manual mechanism used to compress the specimens at the aging temperature was neither smooth nor repeatable which affected the peak load value. To better understand the effect of the mode of initial compression on the measured load, and possibly better estimate “equilibrium” physical relaxation times, systematic stress relaxation experiments were performed using an Instron machine with a thermal chamber. At a given temperature, the DIW polymer was compressed to a fixed strain in either a stepped or continuous manner at two different rates, then held at that strain for 24 hrs. The results indicated that, at a given temperature, the different stress relaxation curves appeared to converge to the same curve at some “equilibrium” time when the non-age-related physical relaxation was mostly complete. Though this observation suggests that the discontinuous manual compression employed by the vessel is feasible, a compression mechanism that is rapid, smooth, and repeatable would enhance its use.

36 MATERIALS SCIENCE↗

Knowledge Distillation for Anomaly Detection

Unsupervised deep learning techniques are widely used to identify anomalous behaviour. The performance of such methods is a product of the amount of training data and the model size. However, the size is often a limiting factor for the deployment on resource-constrained devices. Here, we present a novel procedure based on knowledge distillation for compressing an unsupervised anomaly detection model into a supervised deployable one and we suggest a set of techniques to improve the detection sensitivity. Compressed models perform comparably to their larger counterparts while significantly reducing the size and memory footprint.

Pol, Adrian Alan↗

Robustness of deep learning algorithms in astronomy -- galaxy morphology studies

Deep learning models are being increasingly adopted in wide array of scientific domains, especially to handle high-dimensionality and volume of the scientific data. However, these models tend to be brittle due to their complexity and overparametrization, especially to the inadvertent adversarial perturbations that can appear due to common image processing such as compression or blurring that are often seen with real scientific data. It is crucial to understand this brittleness and develop models robust to these adversarial perturbations. To this end, we study the effect of observational noise from the exposure time, as well as the worst case scenario of a one-pixel attack as a proxy for compression or telescope errors on performance of ResNet18 trained to distinguish between galaxies of different morphologies in LSST mock data. We also explore how domain adaptation techniques can help improve model robustness in case of this type of naturally occurring attacks and help scientists build more trustworthy and stable models.

79 ASTRONOMY AND ASTROPHYSICS↗

$S$ Hmax orientation in the Alpine region from observations of stress-induced anisotropy of nonlinear elasticity

The orientation of $S$ Hmax is commonly estimated from in situ borehole breakouts and earthquake focal mechanisms. Borehole measurements are expensive, and therefore sparse, and earthquake measurements can only be made in regions with many well-characterized earthquakes. Here, we derive the stress-field orientation using stress-induced anisotropy in nonlinear elasticity. In this method, we measure the strain derivative of velocity as a function of azimuth. We use a natural pump-probe (NPP) approach which consists of measuring elastic wave speed using empirical Green’s functions (probe) at different points of the earth tidal strain cycle (pump). The approach is validated using a larger data set in the Northern Alpine Foreland region where the orientation of maximum horizontal compressive stress is known from borehole breakouts and drilling-induced fractures. The technique resolves NNW-SSW to N-S directed $S$ Hmax which is in good agreement with conventional methods and the recent crustal stress model. We confirm that the NPP method can be applied to dense large-scale seismic arrays. The technique is then applied to the Southern Alps to understand the contemporary stress pattern associated with the ongoing deformation due to counterclockwise rotation of the Adriatic plate with respect to the European plate. Our results explain why the two major faults in Northeastern Italy, the Giudicarie Fault and the Periadriatic Line (Pustertal–Gailtal Fault) are currently inactive, while the currently acting stress field allows faults in Slovenia to deform actively. We have demonstrated that the pump-probe method has the potential to fill in the measurement gap left by conventional approaches, both in terms of regional coverage and in depth.

58 GEOSCIENCES↗