Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An estimate of the momentum deposition in the lower thermosphere by the observed diurnal tide

This paper reports a calculation of the acceleration of the zonal mean flow induced by dissipating tides in the equatorial lower thermosphere. Estimates of the gravest symmetric gravitiational Hough mode (1,1) of the migrating diurnal tide are obtained from monthly composites of global winds observed by the Upper Atmosphere Research Satellite (UARS) High Resolution Doppler Imager (HRDI). Using the principles of classical tidal theory, the tidal momentum flux divergence is computed for a series of monthly mean (1,1) fields from January 1992 to May 1993. The contribution to the mean flow by the leading mode of the migrating tide ranges between -5 and -20 (easterly) m/s/day in the equatorial lower thermosphere. A semiannual variation is noted in the tidal amplitudes and the inferred tidal accelerations. These variations are consistent with observed trends in the zonal mean flow of the lower thermosphere.

Lieberman, Ruth S.↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

Electronic structure prediction of medium and high entropy alloys across composition space

We propose machine learning (ML) models to predict the electron density — the fundamental unknown of a material’s ground state — across the composition space of concentrated alloys. From this, other physical properties can be inferred, enabling accelerated exploration. A significant challenge is that the number of descriptors and sampled compositions required for accurate prediction grows rapidly with species. To address this, we employ Bayesian Active Learning (AL), which minimizes training data requirements by leveraging uncertainty quantification capabilities of Bayesian Neural Networks. Compared to the strategic tessellation of the composition space, Bayesian-AL reduces the number of training data points by a factor of 2.5 for ternary (SiGeSn) and 1.7 for quaternary (CrFeCoNi) systems. We also introduce easy-to-optimize, body-attached-frame descriptors, which respect physical symmetries while keeping descriptor-vector size nearly constant as alloy complexity increases. Our ML models demonstrate high accuracy and generalizability in predicting both electron density and energy across composition space.

materials science↗

Learning to Rank in the Age of Muppets: Effectiveness–Efficiency Tradeoffs in Multi-Stage Ranking

It is well known that rerankers built on pretrained transformer models such as BERT have dramatically improved retrieval effectiveness in many tasks. However, these gains have come at substantial costs in terms of efficiency, as noted by many researchers. In this work, we show that it is possible to retain the benefits of transformer-based rerankers in a multi-stage reranking pipeline by first using feature-based learning-to-rank techniques to reduce the number of candidate documents under consideration without adversely affecting their quality in terms of recall. Applied to the MS MARCO passage and document ranking tasks, we are able to achieve the same level of effectiveness, but with up to 18 increase in efficiency. Furthermore, our techniques are orthogonal to other methods focused on accelerating transformer inference, and thus can be combined for even greater efficiency gains. A higher-level message from our work is that, even though pretrained transformers dominate the modern IR landscape, there are still important roles for “traditional” LTR techniques, and that we should not forget history.

catalysis↗

Searching for Strongly Coupled Dark Sectors with Unsupervised and Generative Learning

Recipient of the URA Early Career Award for groundbreaking searches for dark matter arising from strongly coupled dark sectors with the CMS detector, pioneering work in ML-based model-independent anomaly detection for collider and astrophysics experiments, and leadership in the development of new AI/ML techniques to improve event reconstruction and detector simulation in particle physics, as well as novel strategies to accelerate AI inference and throughput with heterogeneous computing using coprocessors as a service.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

$\mathrm{SageNet}$: Fast Neural Network Emulation of the Stiff-amplified Gravitational Waves from Inflation

Accurate modeling of the inflationary gravitational waves (GWs) requires time-consuming, iterative numerical integrations of differential equations to take into account their backreaction on the expansion history. To improve computational efficiency while preserving accuracy, we present the Stiff-amplified Gravitational-wave Emulator Network (SageNet), a deep learning framework designed to replace conventional numerical solvers (code available at https://github.com/YifangLuo/SageNet). SageNet employs a long short-term memory architecture to emulate the present-day energy density spectrum of the inflationary GWs with possible stiff amplification, Ω GW (f). Trained on a data set of 25,689 numerically generated solutions, SageNet allows accurate reconstructions of Ω GW (f) and generalizes well to a wide range of cosmological parameters; 90.9% of the test emulations with randomly distributed parameters exhibit errors of under 4%. In addition, SageNet demonstrates its ability to learn and reproduce the artificial, adaptive sampling patterns in numerical calculations, which implement denser sampling of frequencies around changes in spectral indices in Ω GW (f). The dual capability of learning both physical and artificial features of the numerical GW spectra establishes SageNet as a robust alternative to exact numerical methods. Finally, our benchmark tests show that SageNet reduces the computation time from tens of seconds to milliseconds, achieving a speedup of ∼10 4 times over standard CPU-based numerical solvers with the potential for further acceleration on GPU hardware. These capabilities make SageNet a powerful tool for accelerating Bayesian inference procedures for extended cosmological models. In a broad sense, the SageNet framework offers a fast, accurate, and generalizable solution to modeling cosmological observables whose theoretical predictions demand costly differential equation solvers.

Astronomy data modeling↗

Graph Neural Network-based Tracking as a Service

Recent studies have shown promising results for track finding in dense environments using Graph Neural Network (GNN)-based algorithms. However, GNN-based track finding is computationally slow on CPUs, necessitating the use of coprocessors to accelerate the inference time. Additionally, the large input graph size demands a large device memory for efficient computation, a requirement not met by all computing facilities used for particle physics experiments, particularly those lacking advanced GPUs. Furthermore, deploying the GNN-based track-finding algorithm in a production environment requires the installation of all dependent software packages, exclusively utilized by this algorithm. These computing challenges must be addressed for the successful implementation of GNN-based track-finding algorithm into production settings. In response, we introduce a ``GNN-based tracking as a service'' approach, incorporating a custom backend within the NVIDIA Triton inference server to facilitate GNN-based tracking. This paper presents the performance of this approach using the Perlmutter supercomputer at NERSC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating Hamiltonian Monte Carlo for Bayesian inference in neural networks and neural operators

Hamiltonian Monte Carlo (HMC) is a powerful and accurate method to sample from the posterior distribution in Bayesian inference. However, HMC techniques are computationally demanding for Bayesian neural networks due to the high dimensionality of the network’s parameter space and the non-convexity of their posterior distributions. Therefore, various approximation techniques, such as variational inference (VI) or stochastic gradient MCMC, are often employed to infer the posterior distribution of the network parameters. Such approximations introduce inaccuracies in the inferred distributions, resulting in unreliable uncertainty estimates. In this work, we propose a hybrid approach that combines inexpensive VI and accurate HMC methods to efficiently and accurately quantify uncertainties in neural networks and neural operators. The proposed approach leverages an initial VI training on the full network. We examine the influence of individual parameters on the prediction uncertainty, which shows that a large proportion of the parameters do not contribute substantially to uncertainty in the network predictions. This information is then used to significantly reduce the dimension of the parameter space, and HMC is performed only for the subset of network parameters that strongly influence prediction uncertainties. This yields a framework for accelerating the full batch HMC for posterior inference in neural networks. We demonstrate the efficiency and accuracy of the proposed framework on deep neural networks and operator networks, showing that inference can be performed for large networks with tens to hundreds of thousands of parameters. Finally, we show that this method can effectively learn surrogates for complex physical systems by modeling the operator that maps from upstream conditions to wall-pressure data on a cone in hypersonic flow.

Bayesian inference↗

Latitudinal variation of speed and mass flux in the acceleration region of the solar wind inferred from spectral broadening measurements

Spectral broadening measurements conducted at S-band (13-cm wavelength) during solar minimum conditions in the heliocentric distance range of 3-8 R(sub O) by Mariner 4, Pioneer 10, Mariner 10, Helios 1, Helios 2, and Viking have been combined to reveal a factor of 2.6 reduction in bandwidth from equator to pole. Since spectral broadening bandwidth depends on electron density fluctuation and solar wind speed, and latitudinal variation of the former is available from coherence bandwidth measurements, the remote sensing spectral broadening measurements provide the first determination of the latitudinal variation of solar wind speed in the acceleration region. When combined with electron density measurements deduced from white-light coronagraphs, this result also leads to the first determination of the latitudinal variation of mass flux in the acceleration region. From equator to pole, solar wind speed increases by a factor of 2.2, while mass flux decreases by a factor of 2.3. These results are consistent with measurements of solar wind speed by multi-station intensity scintillation measurements, as well as measurements of mass flux inferred from Lyman alpha observations, both of which pertain to the solar wind beyond 0.5 AU. The spectral broadening observations, therefore, strengthen earlier conclusions about the latitudinal variation of solar wind speed and mass flux, and reinforce current solar coronal models and their implications for solar wind acceleration and solar wind modeling.

Woo, Richard↗

A GPU-Accelerated Population Generation, Sorting, and Mutation Kernel for an Optimization-Based Causal Inference Model

We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.

Cho, Wendy K. Tam↗

LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.

Chitty-Venkata, Krishna Teja↗

On the High- and Low- Altitude Limits of the Auroral Electric Field Region

Using measurements from the High Altitude Plasma Instrument (HAPI) on the Dynamics-Explorer 1 (DE-1) spacecraft and the Low Altitude Plasma Instrument (LAPI) on Dynamics Explorer 2 (DE 2), we investigate both die high altitude and low altitude extents of the auroral acceleration region. To infer the high altitude limit, we searched the HAPI data base for evidence of upward-directed auroral electric fields located above the spacecraft when the HAPI spacecraft is above 9000 km altitude. We find that such acceleration is common when DE-1 flies through die auroral oval at an altitude of 9,000-11,000 km. At altitudes above 11,000 km, the fraction of the orbits with evidence of at least a 1000 V potential drop above the spacecraft falls, becoming essentially zero above an altitude of 15,000 km. Above that altitude, small (100 V) potential drops are frequently observed, but only rarely are approx. 1 kV potentials observed, typically associated with polar cap or 'theta' arcs or westward traveling surges. To investigate the low-altitude limit of the auroral acceleration region, we use conjunctions of DE 1 and DE 2 along auroral field lines and match the upgoing fluxes of ionospheric ions observed by DE 2 with the flux of accelerated upgoing ions observed at DE 1. Calculating the ionospheric scale height from the ion and electron temperatures and assuming that the parallel flow velocity is independent of height above 800 km, we calculate the altitude at which the upwelling ionospheric ions are effectively completely lost to upward acceleration. The initial lowest-altitude acceleration process could be either a perpendicular acceleration or a parallel electric field, but it must be sufficient to give the entire distribution escape energy. We find that in the two cases studied, near the region of peak auroral potential drop the altitude of this acceleration was around 1700 km (near the O/H neutral crossover altitude), but was significantly higher (approx. 2000 km) near the edges of the arc, where the potential was lower. The composition of the upgoing ion beam was consistent with these heights, being predominately H(+) near the edges and O(+) near the peak.

Reiff, P. H.↗

Vestibular compensation and orientation during locomotion

Body, head, and eye movements were studied in three dimensions while walking and turning to determine the role of the vestibular system in directing gaze and maintaining spatial orientation. The body, head, and eyes were represented as three-dimensional coordinate frames, and the movement of these frames was related to a trajectory frame that described the motion of the body on a terrestrial plane. The axis-angle of the body, head, and eye rotation were then compared to the axis-angle of the rotation of the gravitoinertial acceleration (GIA). We inferred the role of the vestibular system during locomotion and the contributions of the VCR and VOR by examining the interrelationship between these coordinate frames. Straight walking induced head and eye rotations in a compensatory manner to the linear accelerations, maintaining head pointing and gaze along the direction of forward motion. Turning generated a combination of compensation and orientation responses. The head leads and steers the turn while the eyes compensate to maintain stable horizontal gaze in space. Saccades shift horizontal gaze as the turn is executed. The head pitches, as during straight walking. It also rolls so that the head tends to align with the orientation of the GIA. Head orientation changes anticipate orientation changes of the GIA. Eye orientation follows the changes in GIA orientation so that the net orientation gaze is closer to the orientation of the GIA. The study indicates that the vestibular system utilizes compensatory and orienting mechanisms to stabilize spatial orientation and gaze during walking and turning.

Non-NASA Center↗

Observations of VHF emissions from 50-mA electron beam injections in the ionosphere that are associated with beam-induced discharges

Results are presented of observations of strong VHF plasma waves with amplitudes in excess of 0.1 mV/m (Hz)1/2 associated with 50 mA electron beam injections in the ionosphere. Data from three swept-frequency receivers carried on two daughter payloads are analyzed to determine the emission spectra of the electron beam for various energies and currents. These results were obtained from the rocket-borne experiment SCEX 3, NASA flight 39.002 UE, launched February 1, 1990. The accelerator payload also carried photometers, which measured luminosity at wavelengths of 391.4 and 380.5 nm. Several times during electron gun activity the measured luminosity increased much faster than proportional to the beam current. This nonlinear increase is evidence that a discharge is occurring in the vicinity of the accelerator payload. It is inferred from the uniform distribution of the luminosity around the accelerator payload that the observed discharge extends significantly outside the beam cylinder.

Goerke, R. T.↗

Measuring Gravitation Using Polarization Spectroscopy

A proposed method of measuring gravitational acceleration would involve the application of polarization spectroscopy to an ultracold, vertically moving cloud of atoms (an atomic fountain). A related proposed method involving measurements of absorption of light pulses like those used in conventional atomic interferometry would yield an estimate of the number of atoms participating in the interferometric interaction. The basis of the first-mentioned proposed method is that the rotation of polarization of light is affected by the acceleration of atoms along the path of propagation of the light. The rotation of polarization is associated with a phase shift: When an atom moving in a laboratory reference interacts with an electromagnetic wave, the energy levels of the atom are Doppler-shifted, relative to where they would be if the atom were stationary. The Doppler shift gives rise to changes in the detuning of the light from the corresponding atomic transitions. This detuning, in turn, causes the electromagnetic wave to undergo a phase shift that can be measured by conventional means. One would infer the gravitational acceleration and/or the gradient of the gravitational acceleration from the phase measurements.

Matsko, Andrey↗

Latitudinal Variation of Solar Wind Speed and Mass Flux in the Acceleration Region of the Solar Wind during Solar Minimum Inferred from Spectral Broadening measurements

In this paper, we use an aggregate of S-band 2.3 GHz (13 cm) spectral broadening observations conducted during solar minimum conditions by the Mariner 4, Pioneer 10, Mariner 10, Helios 1 & 2 and Viking spacecraft to infer the first measurements of the latitudinal variation of solar wind speed and mass flux in the acceleration region of the solar wind at 3-8 R(sub o).

latitudinal variation of solar wind mass flux in a↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗