Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Neural network compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Generic, Sparse Tensor Core for Neural Networks

Sparse neural network attracts more attention for model compression, fast execution, and power reduction. The state-of-the-art designed sparse tensor core for structured and static sparsity, which did not support well for generic or dynamic sparsity. We design a sparse tensor core to support generic sparsity pruning with a novel hybrid and blocked sparse matrix storage format, HB-ELL, which saves computation and storage while keeping the most significant elements, as well as supporting dynamic sparsity for data flow in neural networks. We achieve better performance with preliminary results than the state-of- the-art on an NVIDIA GPU simulator.

Wu, Xiaolong↗

In situ compression artifact removal in scientific data using deep transfer learning and experience replay

The massive amount of data produced during simulation on high-performance computers has grown exponentially over the past decade, exacerbating the need for streaming compression and decompression methods for efficient storage and transfer of this data---key to realizing the full potential of large-scale computational science. Lossy compression approaches such as JPEG when applied to scientific simulation data realized as a stream of images can achieve good compression rates but at the cost of introducing compression artifacts and loss of information. This paper develops a unified framework for in situ compression artifact removal in which the fully convolutional neural network architectures are combined with scalable training, transfer learning, and experience replay to achieve superior accuracy and efficiency while significantly decreasing the storage footprint as compared with the traditional optimization-based approaches. We demonstrate the proposed approach and compare it with compressed sensing postprocessing and other baseline deep learning models using climate simulations and nuclear reactor simulations, both of which are driven by hyperbolic partial differential equations. Our approach when applied to remove the compression artifacts on the JPEG-compressed nuclear reactor simulation data (using a transfer-trained model that was pretrained on the climate simulation data and updated incrementally as the nuclear reactor simulation progressed), achieved a significant improvement---mean peak signal-to-noise ratio of 42.438 as compared with 27.725 obtained with the compressed sensing approach.

97 MATHEMATICS AND COMPUTING↗

Learning high-dimensional parametric maps via reduced basis adaptive residual networks

We propose a scalable framework for the learning of high-dimensional parametric maps via adaptively constructed residual network (ResNet) maps between reduced bases of the inputs and outputs. When just few training data are available, it is beneficial to have a compact parametrization in order to ameliorate the ill-posedness of the neural network training problem. By linearly restricting high-dimensional maps to informed reduced bases of the inputs, one can compress high-dimensional maps in a constructive way that can be used to detect appropriate basis ranks, equipped with rigorous error estimates. A scalable neural network learning framework is thus to learn the nonlinear compressed reduced basis mapping. Unlike the reduced basis construction, however, neural network constructions are not guaranteed to reduce errors by adding representation power, making it difficult to achieve good practical performance. Inspired by recent approximation theory that connects ResNets to sequential minimizing flows, we present an adaptive ResNet construction algorithm. This algorithm allows for depth-wise enrichment of the neural network approximation, in a manner that can achieve good practical performance by first training a shallow network and then adapting. We prove universal approximation of the associated neural network class for $L^2_v$ functions on compact sets. Our overall framework allows for constructive means to detect appropriate breadth and depth, and related compact parametrizations of neural networks, significantly reducing the need for architectural hyperparameter tuning. Numerical experiments for parametric PDE problems and a 3D CFD wing design optimization parametric map demonstrate that the proposed methodology can achieve remarkably high accuracy for limited training data, and outperformed other neural network strategies we compared against.

42 ENGINEERING↗

Hybrid Approaches for Data Reduction of Spatiotemporal Scientific Applications

Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions.

Li, Xiao↗

Scaling kinetic Monte-Carlo simulations of grain growth with combined convolutional and graph neural networks

Graph neural networks (GNN) have emerged as a promising machine learning method for microstructure simulations such as grain growth. However, accurate modeling of realistic grain boundary networks requires large simulation cells, which GNN has difficulty scaling up to. To alleviate the computational costs and memory footprint of GNN, we suggest a hybrid architecture combining a convolutional neural network (CNN) based bijective autoencoder to compress the spatial dimensions, and a GNN that evolves the microstructure in the latent space of reduced spatial sizes. Our results demonstrate that the new design significantly reduces computational costs with using fewer message passing layer (from 12 down to 3) compared with GNN alone. The reduction in computational cost becomes more pronounced as the spatial size increases, indicating strong computational scalability. For the largest mesh evaluated (160 3 ), our method reduces memory usage and runtime in inference by 117× and 115×, respectively, compared with GNN-only baseline. More importantly, it shows higher accuracy and stronger spatiotemporal capability than the GNN-only baseline, especially in long-term testing. Such combination of scalability and accuracy is essential for simulating realistic material microstructures over extended time scales. The improvements can be attributed to the bijective autoencoder’s ability to compress information losslessly from spatial domain into a high dimensional feature space, thereby producing more expressive latent features for the GNN to learn from, while also contributing its own spatiotemporal modeling capability. Training data are generated from stochastic grain growth simulations, providing realistic variability for learning robust microstructure evolution. Comprehensive system validation confirms that the model is accurate, robust, and scalable.

36 MATERIALS SCIENCE↗

Fast 2D Bicephalous Convolutional Autoencoder for Compressing 3D Time Projection Chamber Data

High-energy large-scale particle colliders produce data at high speed in the order of 1 terabytes per second in nuclear physics and petabytes per second in high energy physics. Developing real-time data compression algorithms to reduce such data at high throughput to fit permanent storage has drawn increasing attention. Specifically, at the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider (RHIC), a time projection chamber is used as the main tracking detector, which records particle trajectories in a volume of three-dimensional (3D) cylinder. The resulting data are usually very sparse with occupancy around 10.8%. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. The 3D convolutional neural network (CNN)-based approach, Bicephalous Convolutional Autoencoder (BCAE), outperforms traditional methods both in compression rate and reconstruction accuracy. BCAE can also utilize the computation power of graphical processing units suitable for deployment in a modern heterogeneous highperformance computing environment. This work introduces two BCAE variants: BCAE++ and BCAE-2D. BCAE++ achieves a 15% better compression ratio and a 77% better reconstruction accuracy measured in mean absolute error compared with BCAE. BCAE-2D treats the radial direction as the channel dimension of an image, resulting in a 3× speedup in compression throughput. In addition, we demonstrate an unbalanced autoencoder with a larger decoder can improve reconstruction accuracy without significantly sacrificing throughput. Lastly, we observe both the BCAE++ and BCAE-2D can benefit more from using half-precision mode in throughput (76 - 79% increase) without loss in reconstruction accuracy. The source code and links to data and pretrained models can be found at https://github.com/BNL-DAQ-LDRD/NeuralCompression_v2

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Real-time semantic segmentation on FPGAs for autonomous vehicles with hls4ml

In this paper, we investigate how field programmable gate arrays can serve as hardware accelerators for real-time semantic segmentation tasks relevant for autonomous driving. Considering compressed versions of the ENet convolutional neural network architecture, we demonstrate a fully-on-chip deployment with a latency of 4.9 ms per image, using less than 30% of the available resources on a Xilinx ZCU102 evaluation board. The latency is reduced to 3 ms per image when increasing the batch size to ten, corresponding to the use case where the autonomous vehicle receives inputs from multiple cameras simultaneously. We show, through aggressive filter reduction and heterogeneous quantization-aware training, and an optimized implementation of convolutional layers, that the power consumption and resource utilization can be significantly reduced while maintaining accuracy on the Cityscapes dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Compliant Intramedullary Stems for Joint Reconstruction

The longevity of current joint replacements is limited by aseptic loosening, which is the primary cause of non-infectious failure for hip, knee, and ankle arthroplasty. Aseptic loosening is typically caused either by osteolysis from particulate wear, or by high shear stresses at the bone-implant interface from over-constraint. Our objective was to demonstrate feasibility of a compliant intramedullary stem that eliminates over-constraint without generating particulate wear. The compliant stem is built around a compliant mechanism that permits rotation about a single axis. We first established several models to understand the relationship between mechanism geometry and implant performance under a given angular displacement and compressive load. We then used a neural network to identify a design space of geometries that would support an expected 100-year fatigue life inside the body. We additively manufactured one representative mechanism for each of three anatomic locations, and evaluated these prototypes on a KR-210 robot. The neural network predicts maximum stress and torsional stiffness with 2.69% and 4.08% error respectively, relative to finite element analysis data. We identified feasible design spaces for all three of the anatomic locations. Simulated peak stresses for the three stem prototypes were below the fatigue limit. Benchtop performance of all three prototypes was within design specifications. Our results demonstrate the feasibility of designing patient- and joint-specific compliant stems that address the root causes of aseptic loosening. Guided by these results, we expect the use of compliant intramedullary stems in joint reconstruction technology to increase implant lifetime.

60 APPLIED LIFE SCIENCES↗

Compressive Neural Representations of Volumetric Scalar Fields

Here, we present an approach for compressing volumetric scalar fields using implicit neural representations. Our approach represents a scalar field as a learned function, wherein a neural network maps a point in the domain to an output scalar value. By setting the number of weights of the neural network to be smaller than the input size, we achieve compressed representations of scalar fields, thus framing compression as a type of function approximation. Combined with carefully quantizing network weights, we show that this approach yields highly compact representations that outperform state-of-the-art volume compression approaches. The conceptual simplicity of our approach enables a number of benefits, such as support for time-varying scalar fields, optimizing to preserve spatial gradients, and random-access field evaluation. We study the impact of network design choices on compression performance, highlighting how simple network architectures are effective for a broad range of volumes.

97 MATHEMATICS AND COMPUTING↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

Zero-Power Analog Optical Processing

The motivation behind this research is the growing challenge of handling the massive amounts of data generated by modern imaging systems. Conventional digital image processing techniques are struggling to keep pace with the demands of high-resolution and high-speed imaging systems for remote sensing due to their high-power consumption and data storage requirements. We present a novel approach based on analog photonics to address this challenge. The proposed system utilizes a silicon-photonics-based image encoder positioned after image formation and initial optical-to-electrical conversion. The photonic encoder compresses image data using a passive disordered photonic structure to perform kernel-type random projections of the raw data. The compressed data is then processed by a back-end neural network, which reconstructs the original image with high fidelity (structural similarity exceeding 90%). Our proposed approach has the potential to compress images with ~ 1000X lower power consumption compared to digital approaches with data rates exceeding 1 terapixel/second.

97 MATHEMATICS AND COMPUTING↗

Deep compressed seismic learning for fast location and moment tensor inferences with natural and induced seismicity

Fast detection and characterization of seismic sources is crucial for decision-making and warning systems that monitor natural and induced seismicity. However, besides the laying out of ever denser monitoring networks of seismic instruments, the incorporation of new sensor technologies such as Distributed Acoustic Sensing (DAS) further challenges our processing capabilities to deliver short turnaround answers from seismic monitoring. In response, this work describes a methodology for the learning of the seismological parameters: location and moment tensor from compressed seismic records. In this method, data dimensionality is reduced by applying a general encoding protocol derived from the principles of compressive sensing. The data in compressed form is then fed directly to a convolutional neural network that outputs fast predictions of the seismic source parameters. Thus, the proposed methodology can not only expedite data transmission from the field to the processing center, but also remove the decompression overhead that would be required for the application of traditional processing methods. An autoencoder is also explored as an equivalent alternative to perform the same job. We observe that the CS-based compression requires only a fraction of the computing power, time, data and expertise required to design and train an autoencoder to perform the same task. Implementation of the CS-method with a continuous flow of data together with generalization of the principles to other applications such as classification are also discussed.

54 ENVIRONMENTAL SCIENCES↗

A Tailored Convolutional Neural Network for Nonlinear Manifold Learning of Computational Physics Data Using Unstructured Spatial Discretizations

In this work, we propose a nonlinear manifold learning technique based on deep convolutional autoencoders that is appropriate for model order reduction of physical systems in complex geometries. Convolutional neural networks have proven to be highly advantageous for compressing data arising from systems demonstrating a slow-decaying Kolmogorov n-width. However, these networks are restricted to data on structured meshes. Unstructured meshes are often required for performing analyses of real systems with complex geometry. Our custom graph convolution operators based on the available differential operators for a given spatial discretization effectively extend the application space of deep convolutional autoencoders to systems with arbitrarily complex geometry that are typically discretized using unstructured meshes. We propose sets of convolution operators based on the spatial derivative operators for the underlying spatial discretization, making the method particularly well suited to data arising from the solution of partial differential equations. We demonstrate the method using examples from heat transfer and fluid mechanics and show better than an order of magnitude improvement in accuracy over linear methods.

97 MATHEMATICS AND COMPUTING↗

Variational encoder geostatistical analysis (VEGAS) with an application to large scale riverine bathymetry

Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow-water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. Further, the reformulation allows variational inference with a small number (e.g., $\mathscr{O}$ (100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.

54 ENVIRONMENTAL SCIENCES↗

Non-intrusive reduced order modeling of natural convection in porous media using convolutional autoencoders: Comparison with linear subspace techniques

Natural convection in porous media is a highly nonlinear multiphysical problem relevant to many engineering applications (e.g., the process of CO 2 sequestration). Here, we extend and present a non-intrusive reduced order model of natural convection in porous media employing deep convolutional autoencoders for the compression and reconstruction and either radial basis function (RBF) interpolation or artificial neural networks (ANNs) for mapping parameters of partial differential equations (PDEs) on the corresponding nonlinear manifolds. To benchmark our approach, we also describe linear compression and reconstruction processes relying on proper orthogonal decomposition (POD) and ANNs. Further, we present comprehensive comparisons among different models through three benchmark problems. The reduced order models, linear and nonlinear approaches, are much faster than the finite element model, obtaining a maximum speed-up of 7 × 10 6 because our framework is not bound by the Courant–Friedrichs–Lewy condition; hence, it could deliver quantities of interest at any given time contrary to the finite element model. Our model’s accuracy still lies within a relative error of 7% in the worst-case scenario. We illustrate that, in specific settings, the nonlinear approach outperforms its linear counterpart and vice versa. We hypothesize that a visual comparison between principal component analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE) could indicate which method will perform better prior to employing any specific compression strategy.

97 MATHEMATICS AND COMPUTING↗

The use of digital thread for reconstruction of local fiber orientation in a compression molded pin bracket via deep learning

A deep convolutional neural network (DCNN) was used for microstructure reconstruction using artificial intelligence (MR-AI) by predicting local average fiber orientation distributions (FOD) in a 3D prepreg platelet molded composite (PPMC) pin bracket. To train the MR-AI model, surface strain fields from residual stresses simulated in PPMC plates were used as the input to the DCNN. A training dataset included PPMC plates with various degrees of global fiber alignment, based on the information obtained from high-fidelity flow simulation of a pin bracket. Further, the MR-AI model was then deployed to analyze FOD in the 3D pin bracket by conducting thermo-elastic residual stress analysis. Initially, the MR-AI model was established entirely on the synthetic simulation data. Then, a μCT scan of a physically molded pin bracket was used to create a finite element model that provided data for additional validation of the DCNN model. For the μCT scan finite element pin bracket the MR-AI model predicted the distribution of fiber orientation tensor components with MAE of 0.10 indicating a global prediction error of 10%. For the flow simulated pin bracket, the MR-AI model predicted the distribution of fiber orientation tensor components with a global prediction error of 11%. The MR-AI model showed the ability to predict regions of varying alignment in the base and flange of the pin bracket. The proposed MR-AI methodology allows for rapid prediction of FOD in geometrically complex parts and offers a promising path to detecting unique fiber orientation states in molded components.

42 ENGINEERING↗

Lossy compression of statistical data using quantum annealer

Abstract We present a new lossy compression algorithm for statistical floating-point data through a representation learning with binary variables. The algorithm finds a set of basis vectors and their binary coefficients that precisely reconstruct the original data. The optimization for the basis vectors is performed classically, while binary coefficients are retrieved through both simulated and quantum annealing for comparison. A bias correction procedure is also presented to estimate and eliminate the error and bias introduced from the inexact reconstruction of the lossy compression for statistical data analyses. The compression algorithm is demonstrated on two different datasets of lattice quantum chromodynamics simulations. The results obtained using simulated annealing show 3–3.5 times better compression performance than the algorithm based on neural-network autoencoder. Calculations using quantum annealing also show promising results, but performance is limited by the integrated control error of the quantum processing unit, which yields large uncertainties in the biases and coupling parameters. Hardware comparison is further studied between the previous generation D-Wave 2000Q and the current D-Wave Advantage system. Our study shows that the Advantage system is more likely to obtain low-energy solutions for the problems than the 2000Q.

97 MATHEMATICS AND COMPUTING↗

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3↗