Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

FOS: Computer and information sciences↗

Collisionless shock formation and the prompt acceleration of solar flare ions

The formation mechanisms of collisionless shocks in solar flare plasmas are investigated. The priamry flare energy release is assumed to arise in the coronal portion of a flare loop as many small regions or 'hot spots' where the plasma beta locally exceeds unity. One dimensional hybrid numerical simulations show that the expansion of these 'hot spots' in a direction either perpendicular or oblique to the ambient magnetic field gives rise to collisionless shocks in a few Omega(i), where Omega(i) is the local ion cyclotron frequency. For solar parameters, this is less than 1 second. The local shocks are then subsequently able to accelerate particles to 10 MeV in less than 1 second by a combined drift-diffusive process. The formation mechanism may also give rise to energetic ions of 100 keV in the shock vicinity. The presence of these energetic ions is due either to ion heating or ion beam instabilities and they may act as a seed population for further acceleration. The prompt acceleration of ions inferred from the Gamma Ray Spectrometer on the Solar Maximum Mission can thus be explained by this mechanism.

Cargill, P. J.↗

Energetic particle abundances in solar electron events

The results of a comprehensive search of the ISEE 3 energetic particle data for solar electron events with associated increases in elements with atomic number Z = 6 or greater are reported. A sample of 90 such events was obtained. The events support earlier evidence of a bimodal distribution in Fe/O or, more clearly, in Fe/C. Most of the electron events belong to the group that is Fe-rich in comparison with the coronal abundance. The Fe-rich events are frequently also He-3-rich and are associated with type III and type V radio bursts and impulsive solar flares. Fe-poor events are associated with type IV bursts and with interplanetary shocks. With some exceptions, event-to-event enhancements in the heavier elements vary smoothly with Z and with Fe/C. In fact, these variations extend across the full range of events despite inferred differences in acceleration mechanism. The origin of source material in all events appears to be coronal and not photospheric.

Reames, D. V.↗

The source of Jovian auroral hiss observed by Voyager 1

Observations of auroral hiss obtained from the Voyager 1 encounter with Jupiter have been reanalyzed. The Jovian auroral hiss was observed near the inner boundary of the warm Io torus and has a low-frequency cutoff caused by propagation near the resonance cone. A simple ray tracing procedure using an offset tilted dipole of the Jovian magnetic field is used to determine possible source locations. The results obtained are consistent with two sources located symmetrically with respect to the centrifugal equator along an L shell (L approximately = 5.59) that is coincident with the boundary between the hot and cold regions of the Io torus and is located just inward of the ribbon feature observed from Earth. The distance of the sources from the centrifugal equator is approximately 0.58 +/- 0.01 R(sub J). Based on the similarity to terrestrial auroral hiss, the Jovian is auroral hiss is believed to be generated by beams of low energy (approximately tens to thousands of eV) electrons. The low-frequency cutoff of the auroral hiss suggests that the electrons are accelerated near the inferred source region, possibly by parallel electric fields similar to those existing in the terrestrial auroral regions. A field-aligned current is inferred to exist at L shells just inward of the plasma ribbon. A possible mechanism for driving this current is discussed.

Morgan, D. D.↗

Disruption of postural readaptation by inertial stimuli following space flight

Postural instability (relative to pre-flight) has been observed in all shuttle astronauts studied upon return from orbital missions. Postural stability was more closely examined in four shuttle astronaut subjects before and after an 8 day orbital mission. Results of the pre- and post-flight postural stability studies were compared with a larger (n = 34) study of astronauts returning from shuttle missions of similar duration. Results from both studies indicated that inadequate vestibular feedback was the most significant sensory deficit contributing to the postural instability observed post flight. For two of the four IML-1 astronauts, post-flight postural instability and rate of recovery toward their earth-normal performance matched the performance of the larger sample. However, post-flight postural control in one returning astronaut was substantially below mean performance. This individual, who was within normal limits with respect to postural control before the mission, indicated that recovery to pre-flight postural stability was also interrupted by a post-flight pitch plane rotation test. A similar, though less extreme departure from the mean recovery trajectory was present in another astronaut following the same post-flight rotation test. The pitch plane rotation stimuli included otolith stimuli in the form of both transient tangential and constant centripetal linear acceleration components. We inferred from these findings that adaptation on orbit and re-adaptation on earth involved a change in sensorimotor integration of vestibular signals most likely from the otolith organs.

NASA Center JSC↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

Schulte, Jan-Frederik [Purdue U.] (ORCID:000000034↗

GPU coprocessors as a service for deep learning inference in high energy physics

In the next decade, the demands for computing in large scientific experiments are expected to grow tremendously. During the same time period, CPU performance increases will be limited. At the CERN Large Hadron Collider (LHC), these two issues will confront one another as the collider is upgraded for high luminosity running. Alternative processors such as graphics processing units (GPUs) can resolve this confrontation provided that algorithms can be sufficiently accelerated. In many cases, algorithmic speedups are found to be largest through the adoption of deep learning algorithms. We present a comprehensive exploration of the use of GPU-based hardware acceleration for deep learning inference within the data reconstruction workflow of high energy physics. We present several realistic examples and discuss a strategy for the seamless integration of coprocessors so that the LHC can maintain, if not exceed, its current performance throughout its running.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Plasma jet effects on the ionospheric plasma

Heavy ion beams were injected into the ionospheric plasma (experiments ARCS 1 and ARCS 2). In ARCS 1, operation of a 25eV argon ion source, mounted on a plasma diagnostic payload, produced an accelerated electron population; broadband electric field turbulence; large, spin synchronized electric field perturbations; and depletions of thermal ions. In ARCS 2, the ion source was deployed upward along the local magnetic field direction away from the diagnostic payload, and observed effects are contained within several meters of the ion source. However, enhanced wave levels near the LHR frequency are observed at distances up to 1 km, as are the injected ions themselves. A measurement of the dominant wavelength of the enhanced waves is consistent with an inference based upon the accelerated electron population seen in ARCS 1. This electron population is not evident during ARCS 2.

Moore, T. E.↗

Simulating the CMS High Granularity Calorimeter with ML

Detector simulation is a key component of physics analysis and related activities in CMS. In the upcoming High Luminosity LHC era, simulation will be required to use a smaller fraction of computing in order to satisfy resource constraints. At the same time, CMS will be upgraded with the new High Granularity Calorimeter (HGCal), which requires significantly more resources to simulate than the existing CMS calorimeters. This computing challenge motivates the use of generative machine learning models as surrogates to replace full physics-based simulation. We study the application of state-of-the-art diffusion models to simulate particle showers in the CMS HGCal. We will discuss methods to overcome the challenges posed by the high-dimensional, irregular geometry of the HGCal. The quality of the showers produced by the diffusion model will be assessed by comparison to the full GEANT4-based simulation. The increase in simulation throughput will be quantified and methods to accelerate the diffusion model inference will also be discussed.

Amram, Oz↗

Motion at the ionization front in the Orion nebula - A kinematic study of the forbidden O I line

The kinematics of the ionization front in H II regions is investigated via a study of the collisionally excited forbidden O I emission at 630 nm. Slit spectra of 5/km s resolution are used to examine the velocities of forbidden O I across the inner few arcminutes of M42. The mean velocity is significantly more positive than for other ions, which confirms the argument that the ionization front is formed against the parent molecular cloud OMC 1. It is inferred that there is acceleration of material in the ionization front at the parent molecular cloud. A new model for M42 based largely on optical emission-line and radio absorption-line studies is presented. A detailed statistical analysis indicates agreement between the structure function and the Kolmogorov theory for turbulence.

O'Dell, C. R.↗

Antarctic Rebound and the Time-Dependence of the Earth's Shape

Great strides have been made during the past 30 years in refining models of the last global glaciation. The refinements draw upon a vastly expanded relative sea level and sedimentary core record. Furthermore, we now possess a sharpened understanding of the mechanisms that drive climate changes associated with deglaciation. Some 15 years ago, using only 5.5 years of ranging data, analyses of the drift in LAGEOS I node acceleration was used to infer that postglacial rebound was responsible for a secular change in the Earth's ellipsoidal shape (Yoder et al., .1983]. Today there exists a wealth of geodynamics satellite orbit data that constrain the secular time-dependence of the Earth's shape and low order gravity field, which includes mass redistribution from present-day glacier and great ice sheet imbalance and from postglacial rebound. We have shown that an unambiguous determination of the secular variation in the Earth's pear shaped harmonic (l = 3, m = 0) might provide information that bears on the present-day mass balance of Antarctica. This issue is revisited in light of new constraints on glacial loading during the late-Pleistocene and Holocene. An especially critical issue for the interpretation of secular odd degree zonal harmonics, l = 3 to 7, is the timing and magnitude of the deglaciation of Antarctica from Last Glacial Maximum. We explore ways in which the recovery of secular variation in both zonal and non-zonal harmonics for l = 2 through 7 can improve constraints on both rebound and present-day ice sheet balance.

Ivins, Erik R.↗

Deducing Electron Properties from Hard X-Ray Observations

X-radiation from energetic electrons is the prime diagnostic of flare-accelerated electrons. The observed X-ray flux (and polarization state) is fundamentally a convolution of the cross-section for the hard X-ray emission process(es) in question with the electron distribution function, which is in turn a function of energy, direction, spatial location and time. To address the problems of particle propagation and acceleration one needs to infer as much information as possible on this electron distribution function, through a deconvolution of this fundamental relationship. This review presents recent progress toward this goal using spectroscopic, imaging and polarization measurements, primarily from the Reuven Ramaty High Energy Solar Spectroscopic Imager (RHESSI). Previous conclusions regarding the energy, angular (pitch angle) and spatial distributions of energetic electrons in solar flares are critically reviewed. We discuss the role and the observational evidence of several radiation processes: free-free electron-ion, free-free electron-electron, free-bound electron-ion, photoelectric absorption and Compton backscatter (albedo), using both spectroscopic and imaging techniques. This unprecedented quality of data allows for the first time inference of the angular distributions of the X-ray-emitting electrons and improved model-independent inference of electron energy spectra and emission measures of thermal plasma. Moreover, imaging spectroscopy has revealed hitherto unknown details of solar flare morphology and detailed spectroscopy of coronal, footpoint and extended sources in flaring regions. Additional attempts to measure hard X-ray polarization were not sufficient to put constraints on the degree of anisotropy of electrons, but point to the importance of obtaining good quality polarization data in the future.

Kontar, E. P.↗

EdgeCortix SAKURA-I Machine-Learning, PCIe Accelerator SEE Heavy Ion Test Report

To enable autonomy in space, machine-learning and computer vision applications become invaluable for sensor processing. However, these algorithms are computationally complex and unfeasible for many embedded central processing units (CPUs) and usually require external coprocessors, such as graphics processing units (GPUs) or accelerators specific to the application, including application specific integrated circuits (ASICs). In power-constrained systems, GPUs tend to consume more power than is acceptable (>40W), so lower-power accelerators have shown promise to provide the performance needed under spacecraft constraints. For radiation engineers, developing methodologies that can properly test CPUs, GPUs, and accelerators, and enable comparisons between them remains a necessary complication to solve as the devices become more complex. The methodology in this test aims to be a start in developing a baseline single-event effect (SEE) test for client-device machine learning accelerators. This category of devices do not host their own operating system. This testing campaign is a continuation of a previous 200 MeV proton test performed in January 2024. This report covers two heavy ion tests of the SAKURA-I card: one in April 2024, and one in June 2024. Additional data was needed after the April test due to ion-range issues experienced at higher linear-energy transfers (LETs). These range issues are described in more detail in Section 8. This experiment characterizes SEEs and data error susceptibility of the EdgeCortix SAKURA-I machine-learning accelerator under heavy ions. The device was monitored for single event upsets (SEUs) and single event functional interrupts (SEFIs) at the Lawrence Berkeley National Laboratory’s 88-inch cyclotron. The SAKURA-I board accelerates machine-learning inference applications on a host computer through a PCIex16 connection. For the purposes of devising an end to end automated analysis workflow for this experiment, the YOLO-V5 and SSD300 objection-detection models, and the ResNet-50, EfficientNet, and MobileNetV2 image classification models were used as a representative suite of analytical machine-learning models.

Seth S Roffe↗

Particle acceleration by propagating interplanetary shocks

The process of charged particle acceleration in interplanetary shocks has been simulated for typical parameters. Since space probe observations of charged particle fluxes are the principle means of inferring the action of an acceleration mechanism, the simulation was designed to follow particles backwards in time from a given observing point, usually 1 AU. This allowed a full diagnosis of which observed particles had interacted with an oncoming shock, where the interacting occurred, and by how much the energy had been changed by the shock interaction. Simple assumptions about the preshock energy spectrum allow the construction of a full temporal profile of the expected intensity, anisotropy, and energy spectrum. The simulations are apparently capable of reproducing the main features of the less than 10 MeV/nuc ion flux enhancements observed at interplanetary shocks using only a laminar interplanetary magnetic field, Archimidean spiral, and a spherical oblique shock with typical speed, strength, and shock normal-magnetic field angle.

Armstrong, T. P.↗

Adapting In Situ Accelerators for Sparsity With Granular Matrix Reordering

Neural network (NN) inference is an essential part of modern systems and is found at the heart of numerous applications ranging from image recognition to natural language processing. In situ NN accelerators can efficiently perform NN inference using resistive crossbars, which makes them a promising solution to the data movement challenges faced by conventional architectures. Although such accelerators demonstrate significant potential for dense NNs, they often do not benefit from sparse NNs, which contain relatively few non-zero weights. Processing sparse NNs on in situ accelerators results in wasted energy to charge the entire crossbar where most elements are zeros. To address this limitation, this paper proposes Granular Matrix Reordering (GMR): a preprocessing technique that enables an energy-efficient computation of sparse NNs on in situ accelerators. GMR reorders the rows and columns of sparse weight matrices to maximize the crossbars' utilization and minimize the total number of crossbars needed to be charged. The reordering process does not rely on sparsity patterns and incurs no accuracy loss. Finally, GMR achieves an average of 28% and up to 34% reduction in energy consumption over seven pruned NNs across four different pruning methods and network architectures.

97 MATHEMATICS AND COMPUTING↗