Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Uncertainties in Density and Simulation Parameters for Radiographic Reconstructions Using Machine Learning

We will develop new and improved methods that will enable a fuller and more extensive use of experimental data, such as from radiographic imaging, towards better characterizing and reducing uncertainties in predictive modeling of weapons performance. We will achieve this by leveraging recent developments in the areas of computational imaging, statistical and machine learning, and reduced order modeling. Key new aspects of our approach include (a) a coupling of deep learning based density reconstruction with fast hydrodynamics simulators that will permit the development of certain guarantees on extrapolatability and bounds on certain other hydrodynamics-related para- metric uncertainties and (b) a dual approach towards better characterizing uncertainties in density reconstruction using deep learning based surrogates to accelerate Bayesian reconstructions on the one hand and using variational inference and Bayesian machine learning on the other hand.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Simultaneous measurement of visible energy and momentum transfer in anti-electron neutrino interactions on hydrocarbon

Precise knowledge of neutrino interaction cross sections is required for current and future accelerator-based neutrino oscillation experiments. Precision neutrino oscillation measurements require inference of the neutrino energy and flavor from the visible particles in the neutrino interaction in the detector. This inference is different for the true neutrino flavors measured, electron and muon neutrinos, and can be studied by observation of neutrino interactions in an experiment’s near detector. However, anti-electron neutrinos make up only a few percent of an anti-muon neutrino beam and pose a challenge in making cross section measurements and predictions. The reported measurements were made using data from MINERvA, a neutrino-nucleus scattering experiment, with an anti-neutrino beam configuration of mean energy ~ 6 GeV. This thesis provides two double-differential cross sections of anti-electron neutrino inclusive charged-current reactions using the kinematics of visible energy, three-momentum transfer, and transverse momentum. The analysis is carried out at low three-momentum transfer, making it sensitive to regions with multi-nucleon effects.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

3D-ReG: A 3D ReRAM-based Heterogeneous Architecture for Training Deep Neural Networks

Deep neural network (DNN) models are being expanded to a broader range of applications. The computational capability of traditional hardware platforms cannot accommodate the growth of model complexity. Among recent technologies to accelerate DNN, resistive memory (ReRAM)-based processing-in-memory (PIM) emerged as a promising solution for DNN inference due to its high efficiency for matrix-based computation. We face two major technical challenges in extending the use of ReRAM-based accelerators for training: (1) full-precision data is essential in back-propagation; (2) the need to support both feed-forward and back-propagation aggravates the data-movement burden. We propose a heterogeneous architecture named as 3D-ReG, which leverages full-precision GPU to ensure training accuracy and low-overhead 3D integration to provide low-cost data movements. Moreover, we introduce conservative and aggressive task-mapping schemes, which partition the computation phases in different ways to balance execution efficiency and training accuracy. We evaluate 3D-ReG implemented with two 3D integration technologies, through-silicon vias (TSVs) and monolithic inter-tier vias (MIVs), and compare them with GPU-only and PIM-only counterparts. Various GPU-only platforms using two main-memory technologies (DRAM, ReRAM) and three interconnect technologies (2D, TSV, MIV) are evaluated as well. Experimental results show that 3D-ReG can achieve on average 5.64× training speedup and 3.56× higher energy efficiency compared with the GPU with DRAM as main memory, at the cost of 0.05%–3.39% accuracy drop. We define a new metric, gain-loss ratio (GLR), which quantitatively evaluates the capability of a DNN training hardware in terms of the model accuracy and hardware efficiency. The results of our comparison show that the aggressive task-mapping scheme on MIV-based 3D-ReG outperforms the other methods.

Computer Science↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗

Simulating Global Terrestrial Carbon and Nitrogen Biogeochemical Cycles With Implicit and Explicit Representations of Soil Microbial Activity

Abstract Nutrient limitation is widespread in terrestrial ecosystems. Accordingly, representations of nitrogen (N) limitation in land models typically dampen rates of terrestrial carbon (C) accrual, compared with C‐only simulations. These previous findings, however, rely on soil biogeochemical models that implicitly represent microbial activity and physiology. Here we present results from a biogeochemical model testbed that allows us to investigate how an explicit versus implicit representation of soil microbial activity, as represented in the MIcrobial‐MIneral Carbon Stabilization (MIMICS) and Carnegie‐Ames‐Stanford Approach (CASA) soil biogeochemical models, respectively, influence plant productivity, and terrestrial C and N fluxes at initialization and over the historical period. When forced with common boundary conditions, larger soil C pools simulated by the MIMICS model reflect longer inferred soil organic matter (SOM) turnover times than those simulated by CASA. At steady state, terrestrial ecosystems experience greater N limitation when using the MIMICS‐CN model, which also increases the inferred SOM turnover time. Over the historical period, however, warming‐induced acceleration of SOM decomposition over high latitude ecosystems increases rates of N mineralization in MIMICS‐CN. This reduces N limitation and results in faster rates of vegetation C accrual. Moreover, as SOM stoichiometry is an emergent property of MIMICS‐CN, we highlight opportunities to deepen understanding of sources of persistent SOM and explore its potential sensitivity to environmental change. Our findings underscore the need to improve understanding and representation of plant and microbial resource allocation and competition in land models that represent coupled biogeochemical cycles under global change scenarios.

54 ENVIRONMENTAL SCIENCES↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

DAmodel: hierarchical Bayesian modelling of DA white dwarfs for spectrophotometric calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards (⁠16.5 < V , 19.5⁠) alongside three CALSPEC standards, from 912 Å to 32 μm. The framework is the first of its kind to jointly infer photometric zero points and WD parameters (surface gravity log g⁠, effective temperature T eff ⁠, extinction A V ⁠, dust relation parameter R V ) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy, and ground-based optical spectroscopy to sub-per cent precision. Photometric residuals for the sample are the lowest yet yielding < 0.004 mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory’s Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

methods: statistical↗

GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design

Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art graph learning model. However, it remains notoriously challenging to inference GCNs over large graph datasets, limiting their application to large real-world graphs and hindering the exploration of deeper and more sophisticated GCN graphs. This is because real-world graphs can be extremely large and sparse. Furthermore, the node degree of GCNs tends to follow the power-law distribution and therefore have highly irregular adjacency matrices, resulting in prohibitive inefficiencies in both data processing and movement and thus substantially limiting the achievable GCN acceleration efficiency. To this end, this paper proposes the first GCN algorithm and accelerator Co-Design framework dubbed GCoD which can largely alleviate the aforementioned GCN irregularity and boost GCNs' inference efficiency. Specifically, on the algorithm level, GCoD integrates a divide and conquer GCN training strategy that polarizes the graphs to be either denser or sparser in local neighborhoods without compromising the model accuracy, resulting in graph adjacency matrices that (mostly) have merely two levels of workload and enjoys largely enhanced regularity and thus ease of acceleration. On the hardware level, we further develop a dedicated two-pronged accelerator with a separated engine to process each of the aforementioned workloads, further boosting the overall utilization and acceleration efficiency. Extensive experiments and ablation studies validate that our GCoD consistently outperforms state-of-the-art designs in terms of accelerator efficiency while maintaining or even improving the task accuracy. Additionally, we visualize GCoD trained graph adjacency matrices to better understand its advantages. All codes and pre-trained models will be released upon acceptance.

You, Haoran↗

The heliospheric ambipolar potential inferred from sunward-propagating halo electrons

ABSTRACT We provide evidence that the sunward-propagating half of the solar wind electron halo distribution evolves without scattering in the inner heliosphere. We assume the particles conserve their total energy and magnetic moment, and perform a ‘Liouville mapping’ on electron pitch angle distributions measured by the Parker Solar Probe SPAN-E instrument. Namely, we show that the distributions are consistent with Liouville’s theorem if an appropriate interplanetary potential is chosen. This potential, an outcome of our fitting method, is compared against the radial profiles of proton bulk flow energy. We find that the inferred potential is responsible for nearly 100 per cent of the proton acceleration in the solar wind at heliocentric distances 0.18-0.79 AU. These observations combine to form a coherent physical picture: the same interplanetary potential accounts for the acceleration of the solar wind protons as well as the evolution of the electron halo. In this picture the halo is formed from a sunward-propagating population that originates somewhere in the outer heliosphere by a yet-unknown mechanism.

79 ASTRONOMY AND ASTROPHYSICS↗

Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU System

The deep learning revolution has been enabled in large part by GPUs, and more recently accelerators, which make it possible to carry out computationally demanding training and inference in acceptable times. As the size of machine learning networks and workloads continues to increase, multi-GPU machines have emerged as an important platform offered on High Performance Computing and cloud data centers. Since these machines are shared among multiple users, it becomes increasingly important to protect applications against potential attacks. In this paper, we explore the vulnerability of Nvidia's DGX multi-GPU machines to covert and side channel attacks. These machines consist of a number of discrete GPUs that are interconnected through a combination of custom interconnect (NVLink) and PCIe connections. We reverse engineer the interconnected cache hierarchy and show that it is possible for an attacker on one GPU to cause contention on the L2 cache of another GPU. We use this observation to first develop a covert channel attack across two GPUs, achieving the best bandwidth of around 4 MB/s. We also develop a prime and probe attack on a remote GPU allowing an attacker to recover the cache access pattern of another workload. This access pattern can be used in any number of side channel attacks: we demonstrate a proof of concept attack that fingerprints the application running on the remote GPU, with high accuracy. We also develop a proof of concept attack to extract hyperparameters of a machine learning workload. Our work establishes for the first time the vulnerability of these machines to microarchitectural attacks and can guide future research to improve their security.

Dutta, Sankha↗

Using "AI Poincare" to analyze non-linear integrable optics

This study dives into the applicability of using automated discovery of conserved quantities in dynamical systems relevant to accelerator physics. Specifically, we explore the performance of AI Poincaré in analyzing numerical trajectory data obtained using the McMillan system of non-linear integrable optics. A comprehensive evaluation of the algorithm's performance is conducted through diverse methodologies. These include the analysis of the estimated number of conserved quantities embedded in a dataset and the deviation of interpolated points on the inferred manifold with respect to points in actually in the dataset. the investigation identifies an optimal range of perturbation distances where the underlying manifold extraction algorithm inside AI Poincaré exhibits optimal performance. Additionally, an improved neural network architecture is proposed based on the observed results. Finally, we apply the algorithm to preliminary experimental data from the Integrable Optics Test Accelerator at Fermilab to successfully infer the number of conserved quantities even in the presence of fast decoherence of the measured signal.

Osmanov, Lazare [Free U. Tbilisi]↗

Optimal Power Management for Large-Scale Battery Energy Storage Systems via Bayesian Inference

Large-scale battery energy storage systems (BESS) have found ever-increasing use across industry and society to accelerate clean energy transition and improve energy supply reliability and resilience. However, their optimal power management poses significant challenges: the underlying high-dimensional nonlinear nonconvex optimization lacks computational tractability in real-world implementation, and the uncertainty of the exogenous power demand makes exact optimization difficult. This paper presents a new solution framework to address these bottlenecks. The solution pivots on introducing power-sharing ratios to specify each cell’s power quota from the output power demand. To find the optimal power-sharing ratios, we formulate a nonlinear model predictive control (NMPC) problem to achieve power-loss-minimizing BESS operation while complying with safety, cell balancing, and power supply-demand constraints. We then propose a parameterized control policy for the power-sharing ratios, which utilizes only three parameters, to reduce the computational demand in solving the NMPC problem. This policy parameterization allows us to translate the NMPC problem into a Bayesian inference problem for the sake of 1) computational tractability, and 2) overcoming the nonconvexity of the optimization problem. We leverage the ensemble Kalman inversion technique to solve the parameter estimation problem. Concurrently, a low-level control loop is developed to seamlessly integrate our proposed approach with the BESS to ensure practical implementation. This low-level controller receives the optimal power-sharing ratios, generates output power references for the cells, and maintains a balance between power supply and demand despite uncertainty in output power. We conduct extensive simulations and experiments on a 20-cell prototype to validate the proposed approach.

Battery energy storage systems (BESSs)↗

HD-Bind: Encoding of Molecular Structure with Low Precision, Hyperdimensional Binary Representations

Publicly available collections of drug-like molecules have grown to comprise tens of billions of compounds due to advances in combinatorial chemistry. Traditional methods for identifying "hit" molecules from a large collection of potential drug-like candidates have relied on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have the major drawback that they require exceptional computing capabilities for even relatively small collections of molecules. Hyperdimensional Computing (HDC) is a recently-proposed learning paradigm that represents data with high-dimension binary vectors; this allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas. We consider existing HDC approaches for molecular property classification and introduce two novel encodings of a commonly-used molecular representation, the extended connectivity fingerprint (ECFP). We show that HDC-based inference methods are as much as 91 times more efficient than traditional machine learning methods, and achieve an acceleration of nearly nine orders of magnitude compared to molecular docking. Our results show that HDC accelerated methods retain competitive accuracy on a number of well-studied tasks such as molecular property predictions using the MoleculeNet dataset, and bind/no-bind activity classification using the DUD-E and LIT-PCBA datasets. Our work thus motivates further investigation into molecular representation learning to develop ultraefficient pre-screening tools.

Jones, William↗

H-GCN: A Graph Convolutional Network Accelerator on Versal ACAP Architecture

Recently Graph Neural Networks (GNNs) have drawn tremendous attentions due to their unique capability to extend the Machine Learning (ML) approaches to broadly defined applications with unstructured data, especially graphs. Comparing with other ML modalities, the acceleration of GNNs is as critical but even more challenging due to the irregularity and heterogeneity from graph typologies that together limit the performance. Existing efforts mainly focus on handling graphs’ irregularity, however, have not studied the heterogeneity. To this end, in this work, we propose H-GCN, a PL-AIE-based hybrid accelerator that leverages the emerging heterogeneity of Xilinx Versal ACAPs to achieve high-performance GNN inference. In particular, H-GCN partitions each graph into three subgraphs based on its inherent heterogeneity and processes them using PL and the newly emerged AIE respectively. To further improve the performance, we explore the sparsity support of AIE and develop an efficient density-aware method to map tiles of SpMM onto the systolic tensor array automatically. Compared with the current state-of-the-art GCN accelerator, HGCN achieves on average 1.5× speedups.

Zhang, Chengming↗

Effects of catholyte aging on high-nickel NMC cathodes in sulfide all-solid-state batteries

Sulfide solid-state electrolytes (SSEs) in all-solid-state batteries (SSBs) are recognized for their high ionic conductivity and inherent safety. The LiNi 0.8 Mn 0.1 Co 0.1 O 2 (NMC811) cathode offers a high thermodynamic potential of approximately 3.8 V vs. Li/Li + and a theoretical specific capacity of 200 mA h g −1 . However, the practical utilization of NMC811 in sulfide SSBs faces significant interfacial challenges. The oxidation instability of sulfide solid electrolytes against NMC811 and the formation of the cathode electrolyte interphase (CEI) during cycling lead to degradation and reduced cell performance. Volumetric changes in NMC during lithiation and de-lithiation can also cause detachment from sulfide electrolytes or internal particle cracking. Despite extensive galvanostatic cycling studies to address the issues, the calendar life of sulfide SSBs remains poorly understood. Here, we systematically studied the effects of four different catholytes on the calendar aging of LiNbO 3 (LNO)-coated NMC811, including Li 6 PS 5 Cl (LPSCl), Li 3 InCl 6 –Li 6 PS 5 Cl (LIC–LPSCl), Li 3 YCl 6 –Li 6 PS 5 Cl (LYC–LPSCl), and Li 10 GeP 2 S 12 (LGPS). Our results indicate that LPSCl provides optimal capacity retention when stored at high state-of-charge (SOC) at room temperature, but the LIC–LPSCl cathode shows significant capacity degradation and chemical incompatibility. We also established an effective electrochemical calendar aging testing protocol to simulate daily usage, enabling quick inference of the calendar life of SSBs. In conclusion, this new testing approach accelerates materials selection strategies for high-nickel NMC composite cathodes in sulfide SSBs.

25 ENERGY STORAGE↗

Enhancing Gaussian Process Surrogates for Optimization and Posterior Approximation via Random Exploration

This paper proposes novel noise-free Bayesian optimization strategies that rely on a random exploration step to enhance the accuracy of Gaussian process surrogate models. The new algorithms retain the ease of implementation of the classical GP-UCB algorithm, but the additional random exploration step accelerates their convergence, nearly achieving the optimal convergence rate. Furthermore, to facilitate Bayesian inference with intractable likelihoods, we propose to utilize optimization iterates for maximum a posteriori estimation to build a Gaussian process surrogate model for the unnormalized log-posterior density. We provide bounds for the Hellinger distance between the true and the approximate posterior distributions in terms of the number of design points. We demonstrate the effectiveness of our Bayesian optimization algorithms in nonconvex benchmark objective functions, in a machine learning hyperparameter tuning problem, and in a black-box engineering design problem. The effectiveness of our posterior approximation approach is demonstrated in two Bayesian inference problems for parameters of dynamical systems.

Bayesian inference↗

Machine learning at the Spallation Neutron Source accelerator and target

We describe the ongoing efforts to apply Machine Learning techniques to improve the performance of our accelerator and target. Specially, we are looking to minimize halo beam losses in the absence of a proper physics model, automatically detect and log anomalies in the target support systems such as cooling, and detect and prevent errant beam pulses in the linac. We also describe the infrastructure we use to acquire and stream data to the GPU cluster for training, our code development cycle, and edge computing for model inference. To minimize halo beam losses, we use a Reinforcement Learning technique tested on a virtual accelerator. The target anomaly detection is trained on archived data using incomplete physics models and is made part of the existing target reporting system. The errant beam prevention analyzes beam current and beam phase waveforms as well as accelerator configuration data to predict errant pulses. We also develop continual learning to adapt to changes in the accelerator.

Accelerator Physics↗

Inference as a Service for HEP & NP experiments

This work presents how American Science Cloud resources and services can accelerate scientific discovery across large-scale HEP & NP experiments. Our demonstrators include use cases from intensity and energy frontiers. This is a first end-to-end multi-experiment, multi-facility demonstration of AmSC platforms.

Bhattacharya, M. [Fermilab]↗