Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Deep Learning Based Workflow for Accelerated Industrial X-Ray Computed Tomography

X-ray computed tomography (XCT) is an important tool for high-resolution non-destructive characterization of additively-manufactured metal components. XCT reconstructions of metal components may have beam hardening artifacts such as cupping and streaking which makes reliable detection of flaws and defects challenging. Furthermore, traditional workflows based on using analytic reconstruction algorithms require a large number of projections for accurate characterization - leading to longer measurement times and hindering the adoption of XCT for in-line inspections. In this paper, we introduce a new workflow based on the use of two neural networks to obtain high-quality accelerated reconstructions from sparse-view XCT scans of single material metal parts. The first network, implemented using fully-connected layers, helps reduce the impact of BH in the projection data without the need of any calibration or knowledge of the component material. The second network, a convolutional neural network, maps a low-quality analytic 3D reconstruction to a high-quality reconstruction. Using experimental data, we demonstrate that our method robustly generalizes across several alloys, and for a range of sparsity levels without any need for retraining the networks thereby enabling accurate and fast industrial XCT inspections.

Rahman, Obaid↗

Uncertainty quantification for deep learning in particle accelerator applications

With the advent of increased computational resources and improved algorithms, machine learning-based models are being increasingly applied to complex problems in particle accelerators. However, such data-driven models may provide overly confident predictions with unknown errors and uncertainties. For reliable deployment of machine learning models in high-regret and safety-critical systems such as particle accelerators, estimates of prediction uncertainty are needed along with accurate point predictions. In this investigation, we evaluate Bayesian neural networks (BNN) as an approach that can provide accurate predictions along with reliably quantified uncertainties for particle accelerator problems, and compare their performance with bootstrapped ensembles of neural networks. We select three accelerator setups for this evaluation: a storage ring, a photoinjector, and a linac. The problems span different data volumes and dimensionalities (e.g., scalar predictions as well as image outputs). It is found that BNN provide accurate predictions of the mean along with reliable estimates of predictive uncertainty across the test cases. In this vein, BNN may offer an attractive alternative to deterministic deep learning tools to generate accurate predictions with quantified uncertainties in particle accelerator applications.

43 PARTICLE ACCELERATORS↗

A Study on the Impact of Temperature-Dependent Ferroelectric Switching Behavior in 3D Memory Architecture

The flourishing development of neural networks that require exponentially growing amounts of data has presented an elevated demand for memory footprint. To address this, researchers have been exploring hardware accelerators with innovative memory architectures like 3D memory. These 3D memory architectures offer enhanced storage capacity and processing capabilities, at a cost of rising on-chip temperature during operation. Hafnium Zirconium Oxide (HZO) based Ferroelectric Random Access Memory (FeRAM) is a promising nonvolatile memory candidate in neural network hardware accelerators for its outstanding write performance and reliability. However, its implementation in the architecture regarding the temperature-dependent ferroelectric switching behavior has not been well studied. In this work, we study the thermal impacts on polarization switching through experimental devices and simulation results. We conduct the circuit and architecture-level simulations to showcase that one can exploit this temperature rise to reduce FeRAM's write voltage and write energy due to its unique temperature-activated polarization switching mechanisms. As the on-chip temperature increases to 351K (ambient temperature at 300K) due to neural network workloads, the access energy per bit can be reduced by 27.6% when a dynamic write voltage is applied.

36 MATERIALS SCIENCE↗

Accelerating Finite-temperature Kohn-Sham Density Functional Theory\ with Deep Neural Networks

We present a numerical modeling workflow based on machine learning (ML) which reproduces the the total energies produced by Kohn-Sham density functional theory (DFT) at finite electronic temperature to within chemical accuracy at negligible computational cost. Based on deep neural networks, our workflow yields the local density of states (LDOS) for a given atomic configuration. From the LDOS, spatially-resolved, energy-resolved, and integrated quantities can be calculated, including the DFT total free energy, which serves as the Born-Oppenheimer potential energy surface for the atoms. We demonstrate the efficacy of this approach for both solid and liquid metals and compare results between independent and unified machine-learning models for solid and liquid aluminum. Our machine-learning density functional theory framework opens up the path towards multiscale materials modeling for matter under ambient and extreme conditions at a computational scale and cost that is unattainable with current algorithms.

36 MATERIALS SCIENCE↗

Accelerating Finite-Temperature Kohn-Sham Density Functional Theory with Deep Neural Networks

We present a numerical modeling workflow based on machine learning (ML) which reproduces the total energies produced by Kohn-Sham density functional theory (DFT) at finite electronic temperature to within chemical accuracy at negligible computational cost. Based on deep neural networks, our workflow yields the local density of states (LDOS) for a given atomic configuration. From the LDOS, spatially-resolved, energy-resolved, and integrated quantities can be calculated, including the DFT total free energy, which serves as the Born-Oppenheimer potential energy surface for the atoms. We demonstrate the efficacy of this approach for both solid and liquid metals and compare results between independent and unified machine-learning models for solid and liquid aluminum. Our machine-learning density functional theory framework opens up the path towards multiscale materials modeling for matter under ambient and extreme conditions at a computational scale and cost that is unattainable with current algorithms.

97 MATHEMATICS AND COMPUTING↗

From IMPEL to Impact: Lessons Learned in Accelerating Innovative Building Technologies

The built environment is a complex ecosystem of social institutions and physical infrastructures. Innovation and entrepreneurship in the building industry are critical levers for market transformation toward equitable climate action. However, climate tech innovation for the built environment is not moving fast enough for global needs, and it lacks fundamental diversity, leading to inequitable outcomes. IMPEL (Incubating Market-propelled Entrepreneurial-mindset at the Labs and Beyond) - a U.S. Department of Energy incubator–addresses these critical issues. Over five years, IMPEL has enabled 250 innovators, including 55% women and diverse founders, to accelerate their buildings and clean energy technologies towards market and climate impact. IMPEL provides access to strategic mentoring and coaching, carbon tools training, testbeds, and powerful public-private pipelines, including industry demonstrations, non-dilutive grants, and venture capital networks. The IMPEL innovation ecosystem has accelerated the pace of innovation and market adoption of building decarbonization technologies. In this paper, leverage the IMPEL stakeholder ecosystem - from innovators to investors and product industry to policymakers - to analyze the critical barriers to decarbonization still encountered in the building industry. We study the IMPEL approach and highlight lessons learned that benefit young businesses pursuing innovative building and building-edge energy technologies to develop new ideas and products. Finally, we propose a ‘market forming’ framework to improve the quality and efficiency of the entrepreneurial ecosystem in the building industry. This framework could scale vetted technologies and the participation of diverse founders to de-risk the climate tech

Singh, Reshma↗

Investigating Scientific Workload Acceleration using BlueField SmartNICs [Slides]

Modern computing platforms whose workloads generate large amounts of network traffic, such as cloud and HPC systems, often suffer from performance bottlenecks associated with the network interface. In order to alleviate the effects of this obstacle, a new generation of accelerators known as ‘SmartNICs’, which are designed to offload low level networking tasks from the processor into the NIC, have emerged.

42 ENGINEERING↗

Generic Multi-Layer Perceptron Inference Accelerator on FPGA (vneuron) v1.0

We have designed and implemented a neural network inference compute engine (vneuron) that can be deployed in the fabric of any FPGA without using special hardware accelerator primitive. The "vneuron" is purely written in verilog, and supports scalable neural network structure with fully connected layers and ReLU activation ( Multi-Layer Perceptron architecture) with 16 bits of precision. We have demonstrated it on an Xilinx Artix 7 FPGA for a 16-input, 8-output MLP with 3 layer, 1600 parameters. It takes 40 DSP48E and 40 BRAM18, and takes 131 clock cycles for computing (1048 ns when clocked at 125MHz). We include PyTorch quantization from a given floating point model, and provide behavioral verification simulation in the disclosed software package.

Du, Qiang↗

Rapid Identification of X-ray Diffraction Patterns Based on Very Limited Data by Interpretable Convolutional Neural Networks

Large volumes of data from material characterizations call for rapid and automatic data analysis to accelerate materials discovery. Herein, we report a convolutional neural network (CNN) that was trained based on theoretical data and very limited experimental data for fast identification of experimental X-ray diffraction (XRD) patterns of metal–organic frameworks (MOFs). To augment the data for training the model, noise was extracted from experimental data and shuffled; then it was merged with the main peaks that were extracted from theoretical spectra to synthesize new spectra. For the first time, one-to-one material identification was achieved. Theoretical MOFs patterns (1012) were augmented to a whole data set of 72 864 samples. It was then randomly shuffled and split into training (58 292 samples) and validation (14 572 samples) data sets at a ratio of 4:1. For the task of discriminating, the optimized model showed the highest identification accuracy of 96.7% for the top 5 ranking on a test data set of 30 hold-out samples. Neighborhood component analysis (NCA) on the experimental XRD samples shows that the samples from the same material are clustered in groups in the NCA map. Analysis on the class activation maps of the last CNN layer further discloses the mechanism by which the CNN model successfully identifies individual MOFs from the XRD patterns. Furthermore, this CNN model trained by the data augmentation technique would not only open numerous potential applications for identifying XRD patterns for different materials, but also pave avenues to autonomously analyze data by other characterization tools such as FTIR, Raman, and NMR spectroscopies.

36 MATERIALS SCIENCE↗

B-DeepONet: An enhanced Bayesian DeepONet for solving noisy parametric PDEs using accelerated replica exchange SGLD

Here, the Deep Operator Network (DeepONet) is a neural network architecture used to approximate operators, including the solution operator of parametric PDEs. DeepONets have shown remarkable approximation ability. However, the performance of DeepONets deteriorates when the training data is polluted with noise, a scenario that occurs in practice. To handle noisy data, we propose a Bayesian DeepONet based on replica exchange Langevin diffusion (reLD). Replica exchange uses two particles. The first particle trains a DeepONet to exploit the loss landscape and make predictions. The other particle trains a different DeepONet to explore the loss landscape and escape local minima via swapping. Compared to DeepONets trained with state-of-the-art gradient-based algorithms (e.g., Adam), the proposed Bayesian DeepONet greatly improves the training convergence for noisy scenarios and accurately estimates the uncertainty. To further reduce the high computational cost of the reLD training of DeepONets, we propose (1) an accelerated training framework that exploits the DeepONet's architecture to reduce its computational cost up to 25% without compromising performance and (2) a transfer learning strategy that accelerates training DeepONets for PDEs with different parameter values. Finally, we illustrate the effectiveness of the proposed Bayesian DeepONet using four parametric PDE problems.

97 MATHEMATICS AND COMPUTING↗

The Integrated Virtual Blast Furnace: Enabling Physics-Based Operational Guidance

As part of a DOE-supported research effort, Purdue University Northwest researchers are collaborating with Oak Ridge National Laboratory and United States Steel Corporation to develop a tool to provide blast furnace operators and engineers with process performance insight comparable to high-fidelity computational fluid dynamics modeling, accelerated to provide “what-if” scenarios at near-real-time speed. This is accomplished by pre-simulating a baseline case and a range of potential operating scenarios to establish how the furnace responds to changing inputs, then training a neural network-based Reduced Order Model to accelerate the speed at which predictions of key parameters can be generated.

Okosun, Tyamo↗

Structure–Property Linkage in Alloys Using Graph Neural Network and Explainable Artificial Intelligence

Deep learning tools have recently shown significant potential for accelerating the prediction of microstructure–property linkage in materials. While deep neural networks like convolution neural networks (CNNs) can extract physics information from 3D microstructure images, they often require a large network architecture and substantial training time. In this research, we trained a graph neural network (GNN) using phase field generated microstructures of Ni-Al alloys to predict the evolution of mechanical properties. We found that a single GNN is capable of accurately predicting the strengthening of Ni-Al alloys with microstructures of varying sizes and dimensions, which cannot otherwise be done with a CNN. Additionally, GNN requires significantly less GPU utilization than CNN and offers more interpretable explanation of predictions using saliency analysis as features are manually defined in the graph. We also utilize explainable artificial intelligence tool Bayesian Inference to determine the coefficients in the power law equation that governs coarsening of precipitates. Overall, our work demonstrates the ability of the GNN to accurately and efficiently extract relevant information from material microstructures without having restrictions on microstructure size or dimension and offers an interpretable explanation.

Chemistry↗

FINETUNA: fine-tuning accelerated molecular simulations

Abstract Progress towards the energy breakthroughs needed to combat climate change can be significantly accelerated through the efficient simulation of atomistic systems. However, simulation techniques based on first principles, such as density functional theory (DFT), are limited in their practical use due to their high computational expense. Machine learning approaches have the potential to approximate DFT in a computationally efficient manner, which could dramatically increase the impact of computational simulations on real-world problems. However, they are limited by their accuracy and the cost of generating labeled data. Here, we present an online active learning framework for accelerating the simulation of atomic systems efficiently and accurately by incorporating prior physical information learned by large-scale pre-trained graph neural network models from the Open Catalyst Project. Accelerating these simulations enables useful data to be generated more cheaply, allowing better models to be trained and more atomistic systems to be screened. We also present a method of comparing local optimization techniques on the basis of both their speed and accuracy. Experiments on 30 benchmark adsorbate-catalyst systems show that our method of transfer learning to incorporate prior information from pre-trained models accelerates simulations by reducing the number of DFT calculations by 91%, while meeting an accuracy threshold of 0.02 eV 93% of the time. Finally, we demonstrate a technique for leveraging the interactive functionality built in to Vienna ab initio Simulation Package (VASP) to efficiently compute single point calculations within our online active learning framework without the significant startup costs. This allows VASP to work in tandem with our framework while requiring 75% fewer self-consistent cycles than conventional single point calculations. The online active learning implementation, and examples using the VASP interactive code, are available in the open source FINETUNA package on Github.

97 MATHEMATICS AND COMPUTING↗

Enforcing constraints for time series prediction in supervised, unsupervised and reinforcement learning

We assume that we are given a time series of data from a dynamical system and our task is to learn the flow map of the dynamical system. We present a collection of results on how to enforce constraints coming from the dynamical system in order to accelerate the training of deep neural networks to represent the flow map of the system as well as increase their predictive ability. In particular, we provide ways to enforce constraints during training for all three major modes of learning, namely supervised, unsupervised and reinforcement learning. In general, the dynamic constraints need to include terms which are analogous to memory terms in model reduction formalisms. Such memory terms act as a restoring force which corrects the errors committed by the learned flow map during prediction. For supervised learning, the constraints are added to the objective function. For the case of unsupervised learning, in particular generative adversarial networks, the constraints are introduced by augmenting the input of the discriminator. Finally, for the case of reinforcement learning and in particular actor-critic methods, the constraints are added to the reward function. In addition, for the reinforcement learning case, we present a novel approach based on homotopy of the action-value function in order to stabilize and accelerate training. We use numerical results for the Lorenz system to illustrate the various constructions.

Stinis, Panagiotis↗

Differentiable programming for online training of a neural artificial viscosity function within a staggered grid Lagrangian hydrodynamics scheme

Lagrangian methods to solve the inviscid Euler equations produce numerical oscillations near shock waves. A common approach to reducing these oscillations is to add artificial viscosity (AV) to the discrete equations. The AV term acts as a dissipative mechanism that attenuates oscillations by smearing the shock across a finite number of computational cells. However, AV introduces several control parameters that are not determined by the underlying physical model, and hence, in practice are tuned to the characteristics of a given problem. We seek to improve the standard quadratic-linear AV form by replacing it with a learned neural function that reduces oscillations relative to exact solutions of the Euler equations, resulting in a hybrid numerical-neural hydrodynamic solver. Because AV is an artificial construct that exists solely to improve the numerical properties of a hydrodynamic code, there is no offline ‘viscosity data’ against which a neural network can be trained before inserting into a numerical simulation, thus requiring online training. We achieve this via differentiable programming, i.e. end-to-end backpropagation or adjoint solution through both the neural and differential equation code, using automatic differentiation of the hybrid code in the Julia programming language to calculate the necessary loss function gradients. A novel offline pre-training step accelerates training by initializing the neural network to the default numerical AV scheme, which can be learned rapidly by space-filling sampling over the AV input space. We find that online training over early time steps of simulation is sufficient to learn a neural AV function that reduces numerical oscillations in long-term hydrodynamic shock simulations. These results offer an early proof-of-principle that online differentiable training of hybrid numerical schemes with novel neural network components can improve certain performance aspects existing in purely numerical schemes.

97 MATHEMATICS AND COMPUTING↗

Machine Learning for Conservative-to-Primitive in Relativistic Hydrodynamics

The numerical solution of relativistic hydrodynamics equations in conservative form requires root-finding algorithms that invert the conservative-to-primitive variables map. These algorithms employ the equation of state of the fluid and can be computationally demanding for applications involving sophisticated microphysics models, such as those required to calculate accurate gravitational wave signals in numerical relativity simulations of binary neutron stars. This work explores the use of machine learning methods to speed up the recovery of primitives in relativistic hydrodynamics. Artificial neural networks are trained to replace either the interpolations of a tabulated equation of state or directly the conservative-to-primitive map. The application of these neural networks to simple benchmark problems shows that both approaches improve over traditional root finders with tabular equation-of-state and multi-dimensional interpolations. In particular, the neural networks for the conservative-to-primitive map accelerate the variable recovery by more than an order of magnitude over standard methods while maintaining accuracy. Neural networks are thus an interesting option to improve the speed and robustness of relativistic hydrodynamics algorithms.

79 ASTRONOMY AND ASTROPHYSICS↗