Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Uncertainty quantification of material parameters in modeling coupled metal and high explosive experiments

Experiments involving the coupling of metal and high explosives (HE) are of notable defense-related interest, and we seek to refine the uncertainty quantification associated with models of such experiments. In particular, our focus is on how uncertainty related to the metal constitutive model challenges our ability to infer high explosive model parameters when analyzing focused science experiments. We consider three focused experiments involving an HE accelerating metal: small plate tests with tantalum/LX-14 and tantalum/LX-17 pairings as well as a tantalum/LX-17 cylinder test. For all three models, we perform sensitivity analysis to ascertain the influence of metal strength on the coupled experimental response. Moreover, we calibrate each model in a Bayesian setting and study the quantification of metal strength on the inference of the HE parameters. Based on our results, we offer guidance for future metal/HE experiments.

36 MATERIALS SCIENCE↗

Neural network accelerator for quantum control

Efficient quantum control is necessary for practical quantum computing implementations with current technologies. Conventional algorithms for determining optimal control parameters are computationally expensive, largely excluding them from use outside of the simulation. Existing hardware solutions structured as lookup tables are imprecise and costly. By designing a machine learning model to approximate the results of traditional tools, a more efficient method can be produced. Such a model can then be synthesized into a hardware accelerator for use in quantum systems. In this study, we demonstrate a machine learning algorithm for predicting optimal pulse parameters. This algorithm is lightweight enough to fit on a low-resource FPGA and perform inference with a latency of 175 ns and pipeline interval of 5 ns with > 0.99 gate fidelity. In the long term, such an accelerator could be used near quantum computing hardware where traditional computers cannot operate, enabling quantum control at a reasonable cost at low latencies without incurring large data bandwidths outside of the cryogenic environment.

43 PARTICLE ACCELERATORS↗

Enabling Precision Neutrino Oscillation Studies with MINERvA

The next generation of neutrino experiments at accelerators aims to establish matter-antimatter asymmetry inneutrinoflavor oscillations. Precision oscillation measurements require inference of neutrino energies and flavor from the products of O(GeV) neutrino interactions on nuclei. I’ll discuss recent results from the MINERvAexperimentand how they help with this inference.

McFarland, Kevin [University of Rochester]↗

A Model of Double Coronal Hard X-Ray Sources in Solar Flares

A number of double coronal X-ray sources have been observed during solar flares by RHESSI, where the two sources reside at different sides of the inferred reconnection site. However, where and how these X-ray-emitting electrons are accelerated remains unclear. Here we present the first model of the double coronal hard X-ray (HXR) sources, where electrons are accelerated by a pair of termination shocks driven by bidirectional fast reconnection outflows. We model the acceleration and transport of electrons in the flare region by numerically solving the Parker transport equation using velocity and magnetic fields from the macroscopic magnetohydrodynamic simulation of a flux rope eruption. We show that electrons can be efficiently accelerated by the termination shocks and high-energy electrons mainly concentrate around the two shocks. The synthetic HXR emission images display two distinct sources extending to >100 keV below and above the reconnection region, with the upper source much fainter than the lower one. The HXR energy spectra of the two coronal sources show similar spectral slopes, consistent with the observations. Our simulation results suggest that the flare termination shock can be a promising particle acceleration mechanism in explaining the double-source nonthermal emissions in solar flares.

79 ASTRONOMY AND ASTROPHYSICS↗

DriveSense: A Noise-Resilient Framework for Driving Mode Identification

Accurate drive mode classification is essential for enhancing the reliability and predictive maintenance of heavy-duty electric trucks. This study proposes a novel fuzzy logic-based framework, DriveSense, for real-time drive mode classification, addressing key challenges such as sensor noise, transitional behaviors, and computational efficiency. The proposed approach integrates a two-stage filtering pipeline, combining adaptive outlier removal and a dynamic Kalman filter to enhance data quality. A fuzzy inference system with smoothened trapezoidal membership functions is then applied to classify driving modes into standstill, constant speed, acceleration, and deceleration while mitigating the effects of noise and edge cases. Performance evaluation using real-world and simulated drive cycles demonstrates significant improvements in classification accuracy (up to 97.8%), F1-score (up to 0.97), and robustness against noise, while reducing false positives. Comparative analysis against baseline models, demonstrates DriveSense’s superior accuracy and generalizability across diverse driving patterns. The framework’s lightweight and interpretable fuzzy inference engine operates with low computational latency, ensuring compatibility with real-time embedded systems typical of heavy-duty electric trucks. Moreover, DriveSense models transitional behaviors through overlapping fuzzy sets and adaptive borderline classification logic, enabling smooth identification of subtle shifts such as rolling stops or gradual deceleration. These results highlight DriveSense’s potential to enhance predictive maintenance strategies, reduce downtime, and support scalable, fleet-wide diagnostics.

Kumar, Praveen [Oak Ridge National Laboratory (ORN↗

Cometary Activity Begins at Kuiper Belt Distances: Evidence from C/2017 K2

We study the development of activity in the incoming long-period comet C/2017 K2 over the heliocentric distance range 9 ≲ r {sub H} ≲ 16 au. The comet continues to be characterized by a coma of submillimeter-sized and larger particles ejected at low velocity. In a fixed co-moving volume around the nucleus we find that the scattering cross section of the coma, C, is related to the heliocentric distance by a power law, C∝r{sub H}{sup -s}, with heliocentric index s = 1.14 ± 0.05. This dependence is significantly weaker than the r {sub H} {sup -2} variation of the insolation as a result of two effects. These are, first, the heliocentric dependence of the dust velocity and, second, a lag effect due to very slow-moving particles ejected long before the observations were taken. A Monte Carlo model of the photometry shows that dust production beginning at r {sub H} ~ 35 au is needed to match the measured heliocentric index, with only a slight dependence on the particle size distribution. Mass-loss rates in dust at 10 au are of order 10{sup 3} kg s{sup -1}, while loss rates in gas may be much smaller, depending on the unknown dust to gas ratio. Consequently, the ratio of the nongravitational acceleration to the local solar gravity, α', may, depending on the nucleus size, attain values of ~10{sup -7} ≲ α' ≲ 10{sup -5}, comparable to values found in short-period comets at much smaller distances. Nongravitational acceleration in C/2017 K2 and similarly distant comets, while presently unmeasured, may limit the accuracy with which we can infer the properties of the Oort cloud from the orbits of long-period comets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Demonstration of plasma mirror capability for the OMEGA Extended Performance laser system

A plasma mirror platform was developed for the OMEGA-EP facility to redirect beams, thus enabling more flexible experimental configurations as well as a platform that can be used in the future to improve laser contrast. The plasma mirror reflected a short pulse focusing beam at 22.5° angle of incidence onto a 12.5 μm thick Cu foil, generating Bremsstrahlung and k α x rays, and accelerating ions and relativistic electrons. By measuring these secondary sources, the plasma mirror key performance metrics of integrated reflectivity and optical quality are inferred. It is shown that for a 5 ± 2 ps, 310 J laser pulse, the plasma mirror integrated reflectivity was 62 ± 13% at an operating fluence of 1670 J cm –2 , and that the resultant short pulse driven particle acceleration and x-ray generation indicate that the on target intensity was 3.1 × 10 18 W cm –2 , which is indicative of a good post-plasma mirror interaction beam optical quality. By deriving the plasma mirror performance metrics from the secondary source scalings, it was simultaneously demonstrated that the plasma mirror is ready for adoption in short pulse particle acceleration and high energy photon generation experiments using the OMEGA-EP system.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Invariant amplitudes, unpolarized cross sections, and polarization asymmetries in neutrino-nucleon and antineutrino-nucleon elastic scattering

At leading order in weak and electromagnetic couplings, cross sections for (anti)neutrino-nucleon elastic scattering are determined by four nucleon form factors that depend on the momentum transfer Q 2 . Including radiative corrections in the Standard Model and potential new physics contributions beyond the Standard Model, eight invariant amplitudes are possible, depending on both Q 2 and the (anti)neutrino energy E ν . We review the definition of these amplitudes and use them to compute both unpolarized and polarized observables including radiative corrections. We show that unpolarized accelerator neutrino cross-section measurements can probe new physics parameter space within the constraints inferred from precision beta decay measurements. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating Radiation Computations for Dynamical Models With Targeted Machine Learning and Code Optimization

Abstract Atmospheric radiation is the main driver of weather and climate, yet due to a complicated absorption spectrum, the precise treatment of radiative transfer in numerical weather and climate models is computationally unfeasible. Radiation parameterizations need to maximize computational efficiency as well as accuracy, and for predicting the future climate many greenhouse gases need to be included. In this work, neural networks (NNs) were developed to replace the gas optics computations in a modern radiation scheme (RTE+RRTMGP) by using carefully constructed models and training data. The NNs, implemented in Fortran and utilizing BLAS for batched inference, are faster by a factor of 1–6, depending on the software and hardware platforms. We combined the accelerated gas optics with a refactored radiative transfer solver, resulting in clear‐sky longwave (shortwave) fluxes being 3.5 (1.8) faster to compute on an Intel platform. The accuracy, evaluated with benchmark line‐by‐line computations across a large range of atmospheric conditions, is very similar to the original scheme with errors in heating rates and top‐of‐atmosphere radiative forcings typically below 0.1 K day −1 and 0.5 W m −2 , respectively. These results show that targeted machine learning, code restructuring techniques, and the use of numerical libraries can yield material gains in efficiency while retaining accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis and mitigation of parasitic resistance effects for analog in-memory neural network acceleration

To support the increasing demands for efficient deep neural network processing, accelerators based on analog in-memory computation of matrix multiplication have recently gained significant attention for reducing the energy of neural network inference. However, analog processing within memory arrays must contend with the issue of parasitic voltage drops across the metal interconnects, which distort the results of the computation and limit the array size. This work analyzes how parasitic resistance affects the end-to-end inference accuracy of state-of-the-art convolutional neural networks, and comprehensively studies how various design decisions at the device, circuit, architecture, and algorithm levels affect the system's sensitivity to parasitic resistance effects. Here, a set of guidelines are provided for how to design analog accelerator hardware that is intrinsically robust to parasitic resistance, without any explicit compensation or re-training of the network parameters.

97 MATHEMATICS AND COMPUTING↗

Whistler Waves Associated With Electron Beams in Magnetopause Reconnection Diffusion Regions

Abstract Whistler waves are often observed in magnetopause reconnection associated with electron beams. We analyze seven MMS crossings surrounding the electron diffusion region (EDR) to study the role of electron beams in whistler excitation. Waves have two major types: (a) Narrow‐band waves with high ellipticities and (b) broad‐band waves that are more electrostatic with significant variations in ellipticities and wave normal angles. While both types of waves are associated with electron beams, the key difference is the anisotropy of the background population, with perpendicular and parallel anisotropies, respectively. The linear instability analysis suggests that the first type of wave is mainly due to the background anisotropy, with the beam contributing additional cyclotron resonance to enhance the wave growth. The second type of broadband waves are excited via Landau resonance, and as seen in one event, the beam anisotropy induces an additional cyclotron mode. The results are supported by particle‐in‐cell simulations. We infer that the first type occurs downstream of the central EDR, where background electrons experience Betatron acceleration to form the perpendicular anisotropy; the second type occurs in the central EDR of guide field reconnection. A parametric study is conducted with linear instability analysis. A beam anisotropy alone of above ∼3 likely excites the cyclotron mode waves. Large beam drifts cause Doppler shifts and may lead to left‐hand polarizations in the ion frame. Future studies are needed to determine whether the observation covers a broader parameter regime and to understand the competition between whistler and other instabilities.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

alternating direction method of multipliers↗

Electron-beam energy reconstruction for neutrino oscillation measurements

Neutrinos exist in one of three types or ‘flavours’—electron, muon and tau neutrinos—and oscillate from one flavour to another when propagating through space. This phenomena is one of the few that cannot be described using the standard model of particle physics (reviewed in ref. 1), and so its experimental study can provide new insight into the nature of our Universe (reviewed in ref. 2). Neutrinos oscillate as a function of their propagation distance (L) divided by their energy (E). Therefore, experiments extract oscillation parameters by measuring their energy distribution at different locations. As accelerator-based oscillation experiments cannot directly measure E, the interpretation of these experiments relies heavily on phenomenological models of neutrino–nucleus interactions to infer E. Here we exploit the similarity of electron–nucleus and neutrino–nucleus interactions, and use electron scattering data with known beam energies to test energy reconstruction methods and interaction models. We find that even in simple interactions where no pions are detected, only a small fraction of events reconstruct to the correct incident energy. More importantly, widely used interaction models reproduce the reconstructed energy distribution only qualitatively and the quality of the reproduction varies strongly with beam energy. This shows both the need and the pathway to improve current models to meet the requirements of next-generation, high-precision experiments such as Hyper-Kamiokande (Japan) and DUNE (USA).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗