Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

alternating direction method of multipliers↗

Electron-beam energy reconstruction for neutrino oscillation measurements

Neutrinos exist in one of three types or ‘flavours’—electron, muon and tau neutrinos—and oscillate from one flavour to another when propagating through space. This phenomena is one of the few that cannot be described using the standard model of particle physics (reviewed in ref. 1), and so its experimental study can provide new insight into the nature of our Universe (reviewed in ref. 2). Neutrinos oscillate as a function of their propagation distance (L) divided by their energy (E). Therefore, experiments extract oscillation parameters by measuring their energy distribution at different locations. As accelerator-based oscillation experiments cannot directly measure E, the interpretation of these experiments relies heavily on phenomenological models of neutrino–nucleus interactions to infer E. Here we exploit the similarity of electron–nucleus and neutrino–nucleus interactions, and use electron scattering data with known beam energies to test energy reconstruction methods and interaction models. We find that even in simple interactions where no pions are detected, only a small fraction of events reconstruct to the correct incident energy. More importantly, widely used interaction models reproduce the reconstructed energy distribution only qualitatively and the quality of the reproduction varies strongly with beam energy. This shows both the need and the pathway to improve current models to meet the requirements of next-generation, high-precision experiments such as Hyper-Kamiokande (Japan) and DUNE (USA).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Microstructural Assessment of Molybdenum Disulfide Coatings Using Nanoindentation Hardness

MoS 2 coatings are used extensively in aerospace and defense applications due to their ultralow friction and high wear resistance. Burnished and resin-bonded MoS 2 coatings are commonly used in these applications due to simplicity in deposition and history of use, despite issues with consistency in coating properties and performance. Physical vapor deposition (PVD) of MoS 2 thin films has emerged as a process alternative in the past 50 years, promising far greater control over film structure and composition but at a greater cost. Despite PVD’s benefits, hesitance to adoption persists in high-consequence applications, not only due to increased costs but variability in resulting coating properties. These variations in properties and subsequent performance are in part due to the complexity of the PVD process and the sensitive interplay between coating process-structure-property relationships. This work aims to demystify the remaining uncertainties of the process-structure-property relationships in PVD MoS 2 . The microstructure and mechanical and tribological properties of 61 different PVD pure MoS 2 coatings are examined herein. Emphasis has been placed on developing performance-based (i.e., hardness, modulus) metrics that can assess microstructural changes (density, orientation, and crystallinity) and be utilized to accelerate process development and coating optimization. Relationships established within suggest that nanoindentation hardness can be used to infer coating performance (i.e., wear rate) and properties (i.e., density, crystalline texture, and stoichiometry). Furthermore, this work demonstrates that PVD MoS 2 coatings close to the theoretical density of MoS 2 consistently have the best tribological performance and can be reliably identified by their hardness.

MoS2↗

Development of a broadband hard x-ray radiography platform for pulsed-power experiments

In this article, we develop and demonstrate a broadband hard x-ray radiography platform at the Zebra Pulsed Power Laboratory that integrates point-projection radiography, bremsstrahlung measurements, and hard x-ray pinhole imaging, designed to diagnose current-driven, cylindrically compressed matter. Initial laser-pulsed-power coupled experiments revealed that intense background radiation generated during 1 MA Zebra current shots overwhelmed laser-produced hard x-rays, obscuring radiographic images. Using combined spectral and spatial diagnostics, we identify energetic electrons accelerated by return currents as the dominant source of background hard x-rays, with electron energies inferred to be 3–4 MeV based on Monte Carlo simulations, and demonstrate mitigation through modifications to the radiation shielding and return-current configuration. The diagnostic platform was validated using a wire-pinch hard x-ray source, allowing radiographs of static 1-mm-diameter aluminum wires to be obtained while simultaneously measuring x-ray source spectra and spatial emission distributions within a single shot. Measured wire transmission profiles were quantitatively reconstructed using radiation transport simulations that incorporate an experimentally inferred two-temperature exponential x-ray spectrum from bremsstrahlung signal analysis and spatially distributed emission sources identified by pinhole imaging. Agreement between measured and simulated transmission profiles demonstrates the validity of the radiographic and x-ray source characterization approach, establishing this diagnostic platform as a promising tool for diagnosing magnetically driven, high-density plasmas relevant to warm dense matter and inertial fusion energy research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Emerging Jets Search, Triton Server Deployment, and Track Quality Development: Machine Learning Applications in High Energy Physics

Machine learning is becoming prevalent in high energy physics, with numerous applications in physics analyses and event reconstruction showing great improvements compared to traditional computing methods. This thesis studies three projects which each propose new avenues for machine learning applications within the high energy physics CMS experiment located at CERN. In the first project, a search for a dark matter signal called “emerging jets” is performed, using graph neural networks to greatly increase sensitivity to the signal’s signature within the data. The result of this dark matter search sets the most stringent exclusion limits to date on theoretical emerging jet models. Motivated by inefficiencies encountered when processing the emerging jet graph neural network at Fermi National Accelerator Laboratory’s computing centers, the second project re-optimizes the computing centers for machine learning inference. This re-optimization uses NVIDIA Triton Inference Servers to process users’ analysis code heterogeneously, therefore achieving high processing throughput and decreasing user time-to-insight. The last project focuses on an upgrade to the CMS experiment’s real-time event selection system which improves physics object reconstruction under harsh processing conditions. A boosted decision tree is used to quickly and efficiently quantify a reconstructed particle’s “track quality” in order to remove particle tracks reconstructed erroneously. In summary, this thesis will not only present examples of how high energy physics can greatly benefit by leveraging machine learning techniques for physics analysis and reconstruction, but will also provide guidance on how the field can prepare for the inevitable increase in machine learning applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗

Beam-ion Studies in NSTX and NSTX-U (Final Technical Report)

The confinement of high energy "fast ions" is crucial for the success of magnetic fusion as a practical energy source. The research performed by UC Irvine on NSTX and NSTX-U provided new information about this important topic in the configuration known as a "spherical tokamak." Injected neutral beams provided the fast ions (also known as "beam ions"). There were five overarching goals of the research. One objective was to measure the confinement of beam ions in order to ascertain if they were accomplishing their desired purpose of transferring their energy to the bulk plasma. A second goal was to better understand instabilities that are driven unstable by the fast ions. A related goal was to use the similarities and differences between NSTX and the DIII-D conventional tokamak to better understand the fast-ion driven instabilities. A fourth goal was to develop new instruments to measure fast ions and techniques that facilitate analysis of the data. The fifth objective was to measure and understand acceleration of beam ions by RF waves. Much was accomplished in all of these five areas. In the first, it was shown that the large magnetic moment of the beam ions did not harm confinement but instabilities known as the "sawtooth" and "long-lived mode" do. In the second area, much attention was devoted to the fast-ion instabilities known as "Alfvén eigenmodes." In particular, Alfvén eigenmode "avalanches" can cause ~ 30% of the fast ions to be lost in a single explosive burst. The comparative studies with DIII-D showed that an instability discovered on NSTX is also important in conventional tokamaks. In the fourth category, several instruments were developed, a number of effects that complicate interpretation of the data were understood, and progress toward a new method to infer the fast-ion distribution function from the data was made. In the fifth category, the measured profile of accelerated fast ions was initially much broader than theoretical predictions but subsequent improvements in the theoretical modeling achieved good agreement with the data.

43 PARTICLE ACCELERATORS↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Phase Space Reconstruction from Accelerator Beam Measurements Using Neural Networks and Differentiable Simulations

Characterizing the phase space distribution of particle beams in accelerators is a central part of accelerator understanding and performance optimization. However, conventional reconstruction-based techniques either use simplifying assumptions or require specialized diagnostics to infer high-dimensional (> $2D$) beam properties. In this Letter, we introduce a general-purpose algorithm that combines neural networks with differentiable particle tracking to efficiently reconstruct high-dimensional phase space distributions without using specialized beam diagnostics or beam manipulations. Furthermore, we demonstrate that our algorithm accurately reconstructs detailed 4D phase space distributions with corresponding confidence intervals in both simulation and experiment using a single focusing quadrupole and diagnostic screen. This technique allows for the measurement of multiple correlated phase spaces simultaneously, which will enable simplified 6D phase space distribution reconstructions in the future.

47 OTHER INSTRUMENTATION↗

Impact of the magnetic horizon on the interpretation of the Pierre Auger Observatory spectrum and composition data

The flux of ultra-high energy cosmic rays reaching Earth above the ankle energy (5 EeV) can be described as a mixture of nuclei injected by extragalactic sources with very hard spectra and a low rigidity cutoff.Extragalactic magnetic fields existing between the Earth and the closest sources can affect the observed CR spectrum by reducing the flux of low-rigidity particles reaching Earth. We perform a combined fit of the spectrum and distributions of depth of shower maximum measured with the Pierre Auger Observatory including the effect of this magnetic horizon in the propagation of UHECRs in the intergalactic space.We find that, within a specific range of the various experimental and phenomenological systematics, the magnetic horizon effect can be relevant for turbulent magnetic field strengths in the local neighbourhood in which the closest sources lieof order B$_{rms}$ ≃ (50–100) nG (20 Mpc/d$_{s}$)( 100 kpc/L$_{coh}$)$^{1/2}$, with d$_{s}$ the typical intersource separation and L$_{coh}$ the magnetic field coherence length. When this is the case,the inferred slope of the source spectrum becomes softer and can be closer to the expectations of diffusive shock acceleration, i.e., ∝ E$^{-2}$.An additional cosmic-ray population with higher source density and softer spectra, presumably also extragalactic and dominating the cosmic-ray flux at EeV energies, is also required to reproduce the overall spectrum and composition results for all energies down to 0.6 EeV.

79 ASTRONOMY AND ASTROPHYSICS↗

Ion Escape from the Ionosphere of Titan

Ions have been observed to flow away from Titan along its induced magnetic tail by the Plasma Science Instrument (PLS) on Voyager 1 and the Cassini Plasma Spectrometer (CAPS) on Cassini. In both cases, the ions have been inferred to be of ionospheric origin. Recent plasma measurements made at another unmagnetized body, Venus, have also observed similar flow in its magnetic tail. Much earlier, the possibility of such flow was inferred when ionospheric measurements made from the Pioneer Venus Orbiter (PVO) were used to derive upward flow and acceleration of H(+), D(+) and O(+) within the nightside ionosphere of Venus. The measurements revealed that the polarization electric field in the ionosphere produced the principal upward force on these light ions. The resulting vertical flow of H(+) and D(+) was found to be the dominant escape mechanism of hydrogen and deuterium, corresponding to loss rates consistent with large oceans in early Venus. Other electrodynamic forces were unimportant because the plasma beta in the nightside ionosphere of Venus is much greater than one. Although the plasma beta is also greater than one on Titan, ion acceleration is expected to be more complex, especially because the subsolar point and the subflow points can be 180 degrees apart. Following what we learned at Venus, upward acceleration of light ions by the polarization electric field opposing gravity in the ionosphere of Titan will be described. Additional electrodynamic forces resulting from the interaction of Saturn's magnetosphere with Titan's ionosphere will be examined using a recent hybrid model.

Hartle, R.↗

A mechanistic model for creep and thermal aging in Alloy 709

This report describes a physics-based model for creep and thermal aging in Alloy 709. Alloy 709 is an advanced austenitic alloy, targeted for use in future Sodium Fast Reactors (SFRs) and other advanced reactors. The material has superior high temperature properties compared to currently qualified 316 and 304 stainless steels. However, the available creep and thermal aging test database for Alloy 709 is significantly more limited compared to the historical materials. The physics-based model developed here is one way to accelerate the qualification of the material by providing more accurate long-term predictions for creep properties and thermal aging, compared to current empirical time-extrapolate techniques. The crystal plasticity finite element model is used to predict the deformation and failure of alloy 709. The same setup for the CPFE model is used in both the baseline model calibration process and the simulation campaigns for parameter inference. Specific constitutive choices are made for Alloy 709 to capture the primary deformation mechanisms. The dislocation creep formulation developed by Hu and Cocks is extended to account for coupled precipitation formation and the grain boundary cavitation model developed by Sham, Needleman, et al. is used to model grain boundary cavitation-induced failure. A novel update algorithm is proposed to render the semi-discrete constitutive update for the Sham-Needleman model unconditionally stable. A progressive calibration approach is adopted based on the observations that several types of material responses can be effectively decoupled. A surrogate model is trained based on full-fledged CPFE simulations to accelerate the forward model evaluations, and stochastic variational inference (SVI) is used to calibrate the unknown microstructural model parameters. The calibrated mechanistic model is used to predict the long-term creep life of Alloy 709, and the predictions are compared against classical empirical approaches.

36 MATERIALS SCIENCE↗

Mechanism of parallel electric fields inferred from observations

Data from various experiments of the Atmosphere Explorer satellites are analyzed to test their consistency with the model of single-particle linear acceleration through a static parallel potential drop. Theoretical current/voltage and energy flux/voltage relations are applied to the data to predict the densities of source current carriers above the acceleration region for upward Birkeland currents. Comparison of these model predictions with simultaneous electron data indicates that the observed field-aligned potential drops are consistent with direct acceleration in the absence of anomalous resistivity.

Yeh, H.-C.↗

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]↗