Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ionizing radiation effects in SONOS-based neuromorphic inference accelerators

Here, we evaluate the sensitivity of neuromorphic inference accelerators based on Silicon-Oxide-Nitride-Oxide-Silicon (SONOS) charge trap memory arrays to total ionizing dose (TID) effects. Data retention statistics were collected for 16 Mbit of 40 nm SONOS digital memory exposed to ionizing radiation from a Co-60 source, showing good retention of the bits up to the maximum dose of 500 krad(Si). Using this data, we formulate a rate-equation-based model for the TID response of trapped charge carriers in the ONO stack, and predict the effect of TID on intermediate device states between ‘program’ and ‘erase’. This model is then used to simulate arrays of low-power, analog SONOS devices that store 8-bit neural network weights and support in situ matrix-vector multiplication. We evaluate the accuracy of the irradiated SONOS-based inference accelerator on two image recognition tasks – CIFAR-10 and the challenging ImageNet dataset – using state-of-the-art convolutional neural networks, such as ResNet-50. We find that across the datasets and neural networks evaluated, the accelerator tolerates a maximum TID between 10 krad(Si) and 100 krad(Si), with deeper networks being more susceptible to accuracy losses due to TID.

43 PARTICLE ACCELERATORS↗

Generic Multi-Layer Perceptron Inference Accelerator on FPGA (vneuron) v1.0

We have designed and implemented a neural network inference compute engine (vneuron) that can be deployed in the fabric of any FPGA without using special hardware accelerator primitive. The "vneuron" is purely written in verilog, and supports scalable neural network structure with fully connected layers and ReLU activation ( Multi-Layer Perceptron architecture) with 16 bits of precision. We have demonstrated it on an Xilinx Artix 7 FPGA for a 16-input, 8-output MLP with 3 layer, 1600 parameters. It takes 40 DSP48E and 40 BRAM18, and takes 131 clock cycles for computing (1048 ns when clocked at 125MHz). We include PyTorch quantization from a given floating point model, and provide behavioral verification simulation in the disclosed software package.

Du, Qiang↗

Hardware-accelerated inference for real-time gravitational-wave astronomy

The field of transient astronomy has seen a revolution with the first gravitational-wave detections and the arrival of multi-messenger observations they enabled. Transformed by the first detection of binary black hole and binary neutron star mergers, computational demands in gravitational-wave astronomy are expected to grow by at least a factor of two over the next five years as the global network of kilometer-scale interferometers are brought to design sensitivity. With the increase in detector sensitivity, real-time delivery of gravitational-wave alerts will become increasingly important as an enabler of multi-messenger followup. In this work, we report a novel implementation and deployment of deep learning inference for real-time gravitational-wave data denoising and astrophysical source identification. This is accomplished using a generic Inference-as-a-Service model that is capable of adapting to the future needs of gravitational-wave data analysis. Overall, our implementation allows seamless incorporation of hardware accelerators and also enables the use of commercial or private (dedicated) as-a-service computing. Based on our results, we propose a paradigm shift in low-latency and offline computing in gravitational-wave astronomy. Such a shift can address key challenges in peak-usage, scalability and reliability, and provide a data analysis platform particularly optimized for deep learning applications. The achieved sub-millisecond scale latency will also be relevant for any machine learning-based real-time control systems that may be invoked in the operation of near-future and next generation ground-based laser interferometers, as well as the front-end collection, distribution and processing of data from such instruments.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Structurally flexible cloud microphysics, observationally constrained at all scales via ML-accelerated Bayesian inference

We discuss the challenge of developing observationally informed parameterizations of microphysics for use at a hierarchy of modeling scales. Our proposed approach is applicable to any domain that suffers from a two-fold parameterization problem, where physical processes are not resolved at the model scale (the first problem), and those processes are uncertain at any scale (the second problem). For such problems, a physical approach facilitates modeling across scales, as well as systematic observational inference accelerated by machine learning (ML) surrogate models, and quantification of physical uncertainties.

54 ENVIRONMENTAL SCIENCES↗

Accelerating cosmological inference with Gaussian processes and neural networks – an application to LSST Y1 weak lensing and galaxy clustering

ABSTRACT Studying the impact of systematic effects, optimizing survey strategies, assessing tensions between different probes and exploring synergies of different data sets require a large number of simulated likelihood analyses, each of which cost thousands of CPU hours. In this paper, we present a method to accelerate cosmological inference using emulators based on Gaussian process regression and neural networks. We iteratively acquire training samples in regions of high posterior probability which enables accurate emulation of data vectors even in high dimensional parameter spaces. We showcase the performance of our emulator with a simulated 3×2 point analysis of LSST-Y1 with realistic theoretical and systematics modelling. We show that our emulator leads to high-fidelity posterior contours, with an order of magnitude speed-up. Most importantly, the trained emulator can be re-used for extremely fast impact and optimization studies. We demonstrate this feature by studying baryonic physics effects in LSST-Y1 3×2 point analyses where each one of our MCMC runs takes approximately 5 min. This technique enables future cosmological analyses to map out the science return as a function of analysis choices and survey strategy.

Astronomy & Astrophysics↗

Accelerating the Inference of the Exa.TrkX Pipeline

Recently, graph neural networks (GNNs) have been successfully used for a variety of particle reconstruction problems in high energy physics, including particle tracking. The Exa.TrkX pipeline based on GNNs demonstrated promising performance in reconstructing particle tracks in dense environments. It includes five discrete steps: data encoding, graph building, edge filtering, GNN, and track labeling. All steps were written in Python and run on both GPUs and CPUs. In this work, we accelerate the Python implementation of the pipeline through customized and commercial GPU-enabled software libraries, and develop a C++ implementation for inferencing the pipeline. The implementation features an improved, CUDA-enabled fixed-radius nearest neighbor search for graph building and a weakly connected component graph algorithm for track labeling. GNNs and other trained deep learning models are converted to ONNX and inferenced via the ONNX Runtime C++ API. The complete C++ implementation of the pipeline allows integration with existing tracking software. We report the memory usage and average event latency tracking performance of our implementation applied to the TrackML benchmark dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference

Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.

47 OTHER INSTRUMENTATION↗

Accelerating Machine Learning Inference with GPUs in ProtoDUNE Data Processing

Abstract We study the performance of a cloud-based GPU-accelerated inference server to speed up event reconstruction in neutrino data batch jobs. Using detector data from the ProtoDUNE experiment and employing the standard DUNE grid job submission tools, we attempt to reprocess the data by running several thousand concurrent grid jobs, a rate we expect to be typical of current and future neutrino physics experiments. We process most of the dataset with the GPU version of our processing algorithm and the remainder with the CPU version for timing comparisons. We find that a 100-GPU cloud-based server is able to easily meet the processing demand, and that using the GPU version of the event processing algorithm is two times faster than processing these data with the CPU version when comparing to the newest CPUs in our sample. The amount of data transferred to the inference server during the GPU runs can overwhelm even the highest-bandwidth network switches, however, unless care is taken to observe network facility limits or otherwise distribute the jobs to multiple sites. We discuss the lessons learned from this processing campaign and several avenues for future improvements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory

In this work, we demonstrate SONOS (silicon-oxide-nitrideoxide- silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a >10× gain in energy efficiency over state-of-the-art digital and analog inference accelerators.

97 MATHEMATICS AND COMPUTING↗

Inference-Optimized AI and High Performance Computing for Gravitational Wave Detection at Scale

We introduce an ensemble of artificial intelligence models for gravitational wave detection that we trained in the Summit supercomputer using 32 nodes, equivalent to 192 NVIDIA V100 GPUs, within 2 h. Once fully trained, we optimized these models for accelerated inference using NVIDIA TensorRT. We deployed our inference-optimized AI ensemble in the ThetaGPU supercomputer at Argonne Leadership Computer Facility to conduct distributed inference. Using the entire ThetaGPU supercomputer, consisting of 20 nodes each of which has 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, our NVIDIA TensorRT-optimized AI ensemble processed an entire month of advanced LIGO data (including Hanford and Livingston data streams) within 50 s. Our inference-optimized AI ensemble retains the same sensitivity of traditional AI models, namely, it identifies all known binary black hole mergers previously identified in this advanced LIGO dataset and reports no misclassifications, while also providing a 3X inference speedup compared to traditional artificial intelligence models. We used time slides to quantify the performance of our AI ensemble to process up to 5 years worth of advanced LIGO data. In this synthetically enhanced dataset, our AI ensemble reports an average of one misclassification for every month of searched advanced LIGO data. We also present the receiver operating characteristic curve of our AI ensemble using this 5 year long advanced LIGO dataset. This approach provides the required tools to conduct accelerated, AI-driven gravitational wave detection at scale.

97 MATHEMATICS AND COMPUTING↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Electrochemical Random-Access Memory: Progress, Perspectives, and Opportunities

Non-von Neumann computing using neuromorphic systems based on analogue synaptic and neuronal elements has emerged as a potential solution to tackle the growing need for more efficient data processing, but progress toward practical systems has been stymied due to a lack of materials and devices with the appropriate attributes. Recently, solid state electrochemical ion-insertion, also known as electrochemical random access memory (ECRAM) has emerged as a promising approach to realize the needed device characteristics. ECRAM is a three terminal device that operates by tuning electronic conductance in functional materials through solid-state electrochemical redox reactions. This mechanism can be considered as a gate-controlled bulk modulation of dopants and/or phases in the channel. Early work demonstrating that ECRAM can achieve nearly ideal analogue synaptic characteristics has sparked tremendous interest in this approach. More recently, the realization that electrochemical ion insertion can be used to tune the electronic properties of many types of materials including transition metal oxides, layered two-dimensional materials, organic and coordination polymers, and that the changes in conductance can span orders of magnitude has further attracted interest in ECRAM as the basis for analogue synaptic elements for inference accelerators as well as for dynamical devices that can emulate a wide range of neuronal characteristics for implementation in analogue spiking neural networks. At its core, ECRAM shares many fundamental aspects with rechargeable batteries, where ion insertion materials are used extensively for their ability to reversibly store charge and energy. Computing applications, however, present drastically different requirements: systems will require many millions of devices, scaled down to tens of nanometers, all while achieving reliable electronic-state tuning at scaled-up rates and endurances, and with minimal energy dissipation and noise. Further, in this review, we discuss the history, basic concepts, recent progress, as well as the challenges and opportunities for different types of ECRAM, broadly grouped by their primary mobile ionic charge carrier, including Li, protons, and oxygen vacancies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗