Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Observation of the 60Fe Nucleosynthesis-Clock Isotope in Galactic Cosmic Rays

Iron-60 (60Fe) is a radioactive isotope in cosmic rays that serves as a clock to infer an upper limit on the time between nucleosynthesis and acceleration. We have used the ACE-CRIS instrument to collect 3.55 105 iron nuclei, with energies 195 to 500 megaelectron volts per nucleon, of which we identify 15 60Fe nuclei. The 60Fe56Fe source ratio is (7.5 2.9) 105. The detection of supernova-produced 60Fe in cosmic rays implies that the time required for acceleration and transport to Earth does not greatly exceed the 60Fe half-life of 2.6 million years and that the 60Fe source distance does not greatly exceed the distance cosmic rays can diffuse over this time, 1 kiloparsec. A natural place for 60Fe origin is in nearby clusters of massive stars.

Binns, W. R.↗

Simultaneous measurement of visible energy and momentum transfer in anti-electron neutrino interactions on hydrocarbon

Precise knowledge of neutrino interaction cross sections is required for current and future accelerator-based neutrino oscillation experiments. Precision neutrino oscillation measurements require inference of the neutrino energy and flavor from the visible particles in the neutrino interaction in the detector. This inference is different for the true neutrino flavors measured, electron and muon neutrinos, and can be studied by observation of neutrino interactions in an experiment’s near detector. However, anti-electron neutrinos make up only a few percent of an anti-muon neutrino beam and pose a challenge in making cross section measurements and predictions. The reported measurements were made using data from MINERvA, a neutrino-nucleus scattering experiment, with an anti-neutrino beam configuration of mean energy ~ 6 GeV. This thesis provides two double-differential cross sections of anti-electron neutrino inclusive charged-current reactions using the kinematics of visible energy, three-momentum transfer, and transverse momentum. The analysis is carried out at low three-momentum transfer, making it sensitive to regions with multi-nucleon effects.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

3D-ReG: A 3D ReRAM-based Heterogeneous Architecture for Training Deep Neural Networks

Deep neural network (DNN) models are being expanded to a broader range of applications. The computational capability of traditional hardware platforms cannot accommodate the growth of model complexity. Among recent technologies to accelerate DNN, resistive memory (ReRAM)-based processing-in-memory (PIM) emerged as a promising solution for DNN inference due to its high efficiency for matrix-based computation. We face two major technical challenges in extending the use of ReRAM-based accelerators for training: (1) full-precision data is essential in back-propagation; (2) the need to support both feed-forward and back-propagation aggravates the data-movement burden. We propose a heterogeneous architecture named as 3D-ReG, which leverages full-precision GPU to ensure training accuracy and low-overhead 3D integration to provide low-cost data movements. Moreover, we introduce conservative and aggressive task-mapping schemes, which partition the computation phases in different ways to balance execution efficiency and training accuracy. We evaluate 3D-ReG implemented with two 3D integration technologies, through-silicon vias (TSVs) and monolithic inter-tier vias (MIVs), and compare them with GPU-only and PIM-only counterparts. Various GPU-only platforms using two main-memory technologies (DRAM, ReRAM) and three interconnect technologies (2D, TSV, MIV) are evaluated as well. Experimental results show that 3D-ReG can achieve on average 5.64× training speedup and 3.56× higher energy efficiency compared with the GPU with DRAM as main memory, at the cost of 0.05%–3.39% accuracy drop. We define a new metric, gain-loss ratio (GLR), which quantitatively evaluates the capability of a DNN training hardware in terms of the model accuracy and hardware efficiency. The results of our comparison show that the aggressive task-mapping scheme on MIV-based 3D-ReG outperforms the other methods.

Computer Science↗

GHRS Observations of Cool, Low-Gravity Stars: The Outer Atmosphere and Wind of the Nearby K Supergiant Lambda Velorum - 5

UV spectra of lambda Velorum taken with the Goddard High Resolution Spectrograph (GHRS) on the Hubble Space Telescope are used to probe the structure of the outer atmospheric layers and wind and to estimate the mass-loss rate from this K5 lb-II supergiant. VLA radio observations at lambda = 3.6 cm are used to obtain an independent check on the wind velocity and mass-loss rate inferred from the UV observations, Parameters of the chromospheric structure are estimated from measurements of UV line widths, positions, and fluxes and from the UV continuum flux distribution. The ratios of optically thin C II] emission lines indicate a mean chromospheric electron density of log N(sub e) approximately equal 8.9 +/- 0.2 /cc. The profiles of these lines indicate a chromospheric turbulence (v(sub 0) approximately equal 25-36 km/s), which greatly exceeds that seen in either the photosphere or wind. The centroids of optically thin emission lines of Fe II and of the emission wings of self-reversed Fe II lines indicate that they are formed in plasma approximately at rest with respect to the photosphere of the star. This suggests that the acceleration of the wind occurs above the chromospheric regions in which these emission line photons are created. The UV continuum detected by the GHRS clearly traces the mean flux-formation temperature as it increases with height in the chromosphere from a well-defined temperature minimum of 3200 K up to about 4600 K. Emission seen in lines of C III] and Si III] provides evidence of material at higher than chromospheric temperatures in the outer atmosphere of this noncoronal star. The photon-scattering wind produces self-reversals in the strong chromospheric emission lines, which allow us to probe the velocity field of the wind. The velocities to which these self-absorptions extend increase with intrinsic line strength, and thus height in the wind, and therefore directly map the wind acceleration. The width and shape of these self-absorptions reflect a wind turbulence of approximately equal 9-21 km/s. We further characterize the wind by comparing the observations with synthetic profiles generated with the Lamers et al. Sobolev with Exact Integration (SEI) radiative transfer code, assuming simple models of the outer atmospheric structure. These comparisons indicate that the wind in 1994 can be described by a model with a wind acceleration parameter beta approximately 0.9, a terminal velocity of 29-33 km/s, and a mass-loss rate approximately 3 x 10(exp -9) solar M/yr. Modeling of the 3.6 cm radio flux observed in 1997 suggests a more slowly accelerating wind (higher beta) and/or a higher mass-loss rate than inferred from the UV line profiles. These differences may be due to temporal variations in the wind or from limitations in one or both of the models. The discrepancy is currently under investigation.

Carpenter, Kenneth G.↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗

Simulating Global Terrestrial Carbon and Nitrogen Biogeochemical Cycles With Implicit and Explicit Representations of Soil Microbial Activity

Abstract Nutrient limitation is widespread in terrestrial ecosystems. Accordingly, representations of nitrogen (N) limitation in land models typically dampen rates of terrestrial carbon (C) accrual, compared with C‐only simulations. These previous findings, however, rely on soil biogeochemical models that implicitly represent microbial activity and physiology. Here we present results from a biogeochemical model testbed that allows us to investigate how an explicit versus implicit representation of soil microbial activity, as represented in the MIcrobial‐MIneral Carbon Stabilization (MIMICS) and Carnegie‐Ames‐Stanford Approach (CASA) soil biogeochemical models, respectively, influence plant productivity, and terrestrial C and N fluxes at initialization and over the historical period. When forced with common boundary conditions, larger soil C pools simulated by the MIMICS model reflect longer inferred soil organic matter (SOM) turnover times than those simulated by CASA. At steady state, terrestrial ecosystems experience greater N limitation when using the MIMICS‐CN model, which also increases the inferred SOM turnover time. Over the historical period, however, warming‐induced acceleration of SOM decomposition over high latitude ecosystems increases rates of N mineralization in MIMICS‐CN. This reduces N limitation and results in faster rates of vegetation C accrual. Moreover, as SOM stoichiometry is an emergent property of MIMICS‐CN, we highlight opportunities to deepen understanding of sources of persistent SOM and explore its potential sensitivity to environmental change. Our findings underscore the need to improve understanding and representation of plant and microbial resource allocation and competition in land models that represent coupled biogeochemical cycles under global change scenarios.

54 ENVIRONMENTAL SCIENCES↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Dynamics of lunar origin and orbital evolution

The considerable differences in bulk composition of the moon and the earth have led most investigators to favor the capture hypothesis of lunar origin. However, upon closer examination all forms of the hypothesis still seem much less plausible dynamically than formation by accretion, i.e., acquisition of the moon in many small pieces rather than as predominantly one body. Models of accretion do suggest that the proto-lunar matter had a significantly different history from the proto-earth matter. A better understanding of collisions is needed to infer the compositional consequences of this history. Recent work on the acceleration of the moon's orbit exacerbates the time scale problem of orbital evolution. However, it now is much clearer that the locus of tidal dissipation is in the oceans and hence that the solution to the time scale problem lies in differing oceanic configurations in the past.

Kaula, W. M.↗

The quasi-rigid rotation of coronal magnetic fields

Spherical harmonic analysis and numerical simulations are used to study the rotational properties of the coronal magnetic field under the assumption that it can be approximated by a current-free extension of the photospheric field. It is found that the rotation rate in the outer corona is determined, principally, by coronal filtering, the global averages of the photospheric rotation rate, and ongoing source eruptions. The present model is able to account for observationally inferred rotational properties. It is suggested that the coronal rotation rate accelerates gradually due to the equatorward migration of sunspots, and that the 27-day equatorial period is approached toward sunspot minimum as the decaying photospheric flux becomes localized near the equator.

Wang, Y.-M.↗

Radio-Loud Coronal Mass Ejections Without Shocks Near Earth

Type II radio bursts are produced by low energy electrons accelerated in shocks driven by corona) mass ejections (CMEs). One can infer shocks near the Sun, in the Interplanetary medium, and near Earth depending on the wavelength range in which the type II bursts are produced. In fact, type II bursts are good indicators of CMEs that produce solar energetic particles. If the type 11 burst occurs from a source on the Earth-facing side of the solar disk, it is highly likely that a shock arrives at Earth in 2-3 days and hence can be used to predict shock arrival at Earth. However, a significant fraction of CMEs producing type II bursts were not associated shocks at Earth, even though the CMEs originated close to the disk center. There are several reasons for the lack of shock at 1 AU. CMEs originating at large central meridian distances (CMDs) may be driving a shock, but the shock may not be extended sufficiently to reach to the Sun-Earth line. Another possibility is CME cannibalism because of which shocks merge and one observes a single shock at Earth. Finally, the CME-driven shock may become weak and dissipate before reaching 1 AU. We examined a set of 30 type II bursts observed by the Wind/WAVES experiment that had the solar sources very close to the disk center (within a CMD of 15 degrees), but did not have shock at Earth. We find that the near-Sun speeds of the associated CMEs average to approx.600 km/s, only slightly higher than the average speed of CMEs associated with radio-quiet shocks. However, the fraction of halo CMEs is only approx.28%, compared to 40% for radio-quiet shocks and 72% for all radio-loud shocks. We conclude that the disk-center radio loud CMEs with no shocks at 1 AU are generally of lower energy and they drive shocks only close to the Sun.

Gopalswamy, N.↗

Conservative force model performance for TOPEX/Poseidon precision orbit determination

The TOPEX/Poseidon spacecraft was launched on August 10, 1992 to study the Earth's oceans. To achieve maximum benefit from the altimetric data collected, mission requirements dictate that TOPEX/Poseidon's orbit must be computed at an unprecedented level of accuracy. In order to satisfy these requirements, a model which accounts for the satellite's complex geometry, attitude, and surface properties has been developed. This `box-wing' representation treats the spacecraft as the combination of flat plates arranged in the shape of a box and a connecetd solar array. The nonconservative forces acting on each of the eight surfaces are computed independently, yielding vector accelerations which are summed to compute the total aggregate effect on the satellite center-of-mass. Parameters associated with each flat plate were derived from a finite element analysis of the spacecraft. Certain parameters can be inferred from tracking data and have been adjusted to obtain a better representation of the satellite acceleration history. Changes in the nominal mission profile and the presence of an `anomalistic' force have complicated this tuning process. Model performance, parameter sensitivities, and the `anomalistic' force will be discussed.

Marshall, J. Andrew↗

DAmodel: hierarchical Bayesian modelling of DA white dwarfs for spectrophotometric calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards (⁠16.5 < V , 19.5⁠) alongside three CALSPEC standards, from 912 Å to 32 μm. The framework is the first of its kind to jointly infer photometric zero points and WD parameters (surface gravity log g⁠, effective temperature T eff ⁠, extinction A V ⁠, dust relation parameter R V ) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy, and ground-based optical spectroscopy to sub-per cent precision. Photometric residuals for the sample are the lowest yet yielding < 0.004 mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory’s Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

methods: statistical↗

GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design

Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art graph learning model. However, it remains notoriously challenging to inference GCNs over large graph datasets, limiting their application to large real-world graphs and hindering the exploration of deeper and more sophisticated GCN graphs. This is because real-world graphs can be extremely large and sparse. Furthermore, the node degree of GCNs tends to follow the power-law distribution and therefore have highly irregular adjacency matrices, resulting in prohibitive inefficiencies in both data processing and movement and thus substantially limiting the achievable GCN acceleration efficiency. To this end, this paper proposes the first GCN algorithm and accelerator Co-Design framework dubbed GCoD which can largely alleviate the aforementioned GCN irregularity and boost GCNs' inference efficiency. Specifically, on the algorithm level, GCoD integrates a divide and conquer GCN training strategy that polarizes the graphs to be either denser or sparser in local neighborhoods without compromising the model accuracy, resulting in graph adjacency matrices that (mostly) have merely two levels of workload and enjoys largely enhanced regularity and thus ease of acceleration. On the hardware level, we further develop a dedicated two-pronged accelerator with a separated engine to process each of the aforementioned workloads, further boosting the overall utilization and acceleration efficiency. Extensive experiments and ablation studies validate that our GCoD consistently outperforms state-of-the-art designs in terms of accelerator efficiency while maintaining or even improving the task accuracy. Additionally, we visualize GCoD trained graph adjacency matrices to better understand its advantages. All codes and pre-trained models will be released upon acceptance.

You, Haoran↗

Time-Varying Changes in the Simulated Structure of the Brewer-Dobson Circulation

A series of simulations using the NASA Goddard Earth Observing System Chemistry Climate Model are analyzed in order to assess changes in the Brewer-Dobson Circulation (BDC) over the past 55 years. When trends are computed over the past 55 years, the BDC accelerates throughout the stratosphere, consistent with previous modeling results. However, over the second half of the simulations (i.e., since the late 1980s), the model simulates structural changes in the BDC as the temporal evolution of the BDC varies between regions in the stratosphere. In the mid-stratosphere in the midlatitude Northern Hemisphere, the BDC does not accelerate in the ensemble mean of our simulations despite increases in greenhouse gas concentrations and warming sea surface temperatures, and it even decelerates in one ensemble member. This deceleration is reminiscent of changes inferred from satellite instruments and in situ measurements. In contrast, the BDC in the lower stratosphere continues to accelerate. The main forcing agents for the recent slowdown in the mid-stratosphere appear to be declining ozone-depleting substance (ODS) concentrations and the timing of volcanic eruptions. Changes in both mean age of air and the tropical upwelling of the residual circulation indicate a lack of recent acceleration. We therefore clarify that the statement that is often made that climate models simulate a decreasing age throughout the stratosphere only applies over long time periods and is not necessarily the case for the past 25 years, when most tracer measurements were taken.

stratospheric ozone↗

The heliospheric ambipolar potential inferred from sunward-propagating halo electrons

ABSTRACT We provide evidence that the sunward-propagating half of the solar wind electron halo distribution evolves without scattering in the inner heliosphere. We assume the particles conserve their total energy and magnetic moment, and perform a ‘Liouville mapping’ on electron pitch angle distributions measured by the Parker Solar Probe SPAN-E instrument. Namely, we show that the distributions are consistent with Liouville’s theorem if an appropriate interplanetary potential is chosen. This potential, an outcome of our fitting method, is compared against the radial profiles of proton bulk flow energy. We find that the inferred potential is responsible for nearly 100 per cent of the proton acceleration in the solar wind at heliocentric distances 0.18-0.79 AU. These observations combine to form a coherent physical picture: the same interplanetary potential accounts for the acceleration of the solar wind protons as well as the evolution of the electron halo. In this picture the halo is formed from a sunward-propagating population that originates somewhere in the outer heliosphere by a yet-unknown mechanism.

79 ASTRONOMY AND ASTROPHYSICS↗

Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU System

The deep learning revolution has been enabled in large part by GPUs, and more recently accelerators, which make it possible to carry out computationally demanding training and inference in acceptable times. As the size of machine learning networks and workloads continues to increase, multi-GPU machines have emerged as an important platform offered on High Performance Computing and cloud data centers. Since these machines are shared among multiple users, it becomes increasingly important to protect applications against potential attacks. In this paper, we explore the vulnerability of Nvidia's DGX multi-GPU machines to covert and side channel attacks. These machines consist of a number of discrete GPUs that are interconnected through a combination of custom interconnect (NVLink) and PCIe connections. We reverse engineer the interconnected cache hierarchy and show that it is possible for an attacker on one GPU to cause contention on the L2 cache of another GPU. We use this observation to first develop a covert channel attack across two GPUs, achieving the best bandwidth of around 4 MB/s. We also develop a prime and probe attack on a remote GPU allowing an attacker to recover the cache access pattern of another workload. This access pattern can be used in any number of side channel attacks: we demonstrate a proof of concept attack that fingerprints the application running on the remote GPU, with high accuracy. We also develop a proof of concept attack to extract hyperparameters of a machine learning workload. Our work establishes for the first time the vulnerability of these machines to microarchitectural attacks and can guide future research to improve their security.

Dutta, Sankha↗

Using "AI Poincare" to analyze non-linear integrable optics

This study dives into the applicability of using automated discovery of conserved quantities in dynamical systems relevant to accelerator physics. Specifically, we explore the performance of AI Poincaré in analyzing numerical trajectory data obtained using the McMillan system of non-linear integrable optics. A comprehensive evaluation of the algorithm's performance is conducted through diverse methodologies. These include the analysis of the estimated number of conserved quantities embedded in a dataset and the deviation of interpolated points on the inferred manifold with respect to points in actually in the dataset. the investigation identifies an optimal range of perturbation distances where the underlying manifold extraction algorithm inside AI Poincaré exhibits optimal performance. Additionally, an improved neural network architecture is proposed based on the observed results. Finally, we apply the algorithm to preliminary experimental data from the Integrable Optics Test Accelerator at Fermilab to successfully infer the number of conserved quantities even in the presence of fast decoherence of the measured signal.

Osmanov, Lazare [Free U. Tbilisi]↗

Mechanisms of Ionospheric Mass Escape

The dependence of ionospheric O+ escape flux on electromagnetic energy flux and electron precipitation into the ionosphere is derived for a hypothetical ambipolar pick-up process, powered the relative motion of plasmas and neutral upper atmosphere, and by electron precipitation, at heights where the ions are magnetized but influenced by photo-ionization, collisions with gas atoms, ambipolar and centrifugal acceleration. Ion pick-up by the convection electric field produces "ring-beam" or toroidal velocity distributions, as inferred from direct plasma measurements, from observations of the associated waves, and from the spectra of incoherent radar echoes. Ring-beams are unstable to plasma wave growth, resulting in rapid relaxation via transverse velocity diffusion, into transversely accelerated ion populations. Ion escape is substantially facilitated by the ambipolar potential, but is only weakly affected by centrifugal acceleration. If, as cited simulations suggest, ion ring beams relax into non-thermal velocity distributions with characteristic speed equal to the local ion-neutral flow speed, a generalized "Jeans escape" calculation shows that the escape flux of ionospheric O+ increases with Poynting flux and with precipitating electron density in rough agreement with observations.

Moore, T. E.↗