Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “inference accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Microstructural Assessment of Molybdenum Disulfide Coatings Using Nanoindentation Hardness

MoS 2 coatings are used extensively in aerospace and defense applications due to their ultralow friction and high wear resistance. Burnished and resin-bonded MoS 2 coatings are commonly used in these applications due to simplicity in deposition and history of use, despite issues with consistency in coating properties and performance. Physical vapor deposition (PVD) of MoS 2 thin films has emerged as a process alternative in the past 50 years, promising far greater control over film structure and composition but at a greater cost. Despite PVD’s benefits, hesitance to adoption persists in high-consequence applications, not only due to increased costs but variability in resulting coating properties. These variations in properties and subsequent performance are in part due to the complexity of the PVD process and the sensitive interplay between coating process-structure-property relationships. This work aims to demystify the remaining uncertainties of the process-structure-property relationships in PVD MoS 2 . The microstructure and mechanical and tribological properties of 61 different PVD pure MoS 2 coatings are examined herein. Emphasis has been placed on developing performance-based (i.e., hardness, modulus) metrics that can assess microstructural changes (density, orientation, and crystallinity) and be utilized to accelerate process development and coating optimization. Relationships established within suggest that nanoindentation hardness can be used to infer coating performance (i.e., wear rate) and properties (i.e., density, crystalline texture, and stoichiometry). Furthermore, this work demonstrates that PVD MoS 2 coatings close to the theoretical density of MoS 2 consistently have the best tribological performance and can be reliably identified by their hardness.

MoS2↗

Development of a broadband hard x-ray radiography platform for pulsed-power experiments

In this article, we develop and demonstrate a broadband hard x-ray radiography platform at the Zebra Pulsed Power Laboratory that integrates point-projection radiography, bremsstrahlung measurements, and hard x-ray pinhole imaging, designed to diagnose current-driven, cylindrically compressed matter. Initial laser-pulsed-power coupled experiments revealed that intense background radiation generated during 1 MA Zebra current shots overwhelmed laser-produced hard x-rays, obscuring radiographic images. Using combined spectral and spatial diagnostics, we identify energetic electrons accelerated by return currents as the dominant source of background hard x-rays, with electron energies inferred to be 3–4 MeV based on Monte Carlo simulations, and demonstrate mitigation through modifications to the radiation shielding and return-current configuration. The diagnostic platform was validated using a wire-pinch hard x-ray source, allowing radiographs of static 1-mm-diameter aluminum wires to be obtained while simultaneously measuring x-ray source spectra and spatial emission distributions within a single shot. Measured wire transmission profiles were quantitatively reconstructed using radiation transport simulations that incorporate an experimentally inferred two-temperature exponential x-ray spectrum from bremsstrahlung signal analysis and spatially distributed emission sources identified by pinhole imaging. Agreement between measured and simulated transmission profiles demonstrates the validity of the radiographic and x-ray source characterization approach, establishing this diagnostic platform as a promising tool for diagnosing magnetically driven, high-density plasmas relevant to warm dense matter and inertial fusion energy research.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Emerging Jets Search, Triton Server Deployment, and Track Quality Development: Machine Learning Applications in High Energy Physics

Machine learning is becoming prevalent in high energy physics, with numerous applications in physics analyses and event reconstruction showing great improvements compared to traditional computing methods. This thesis studies three projects which each propose new avenues for machine learning applications within the high energy physics CMS experiment located at CERN. In the first project, a search for a dark matter signal called “emerging jets” is performed, using graph neural networks to greatly increase sensitivity to the signal’s signature within the data. The result of this dark matter search sets the most stringent exclusion limits to date on theoretical emerging jet models. Motivated by inefficiencies encountered when processing the emerging jet graph neural network at Fermi National Accelerator Laboratory’s computing centers, the second project re-optimizes the computing centers for machine learning inference. This re-optimization uses NVIDIA Triton Inference Servers to process users’ analysis code heterogeneously, therefore achieving high processing throughput and decreasing user time-to-insight. The last project focuses on an upgrade to the CMS experiment’s real-time event selection system which improves physics object reconstruction under harsh processing conditions. A boosted decision tree is used to quickly and efficiently quantify a reconstructed particle’s “track quality” in order to remove particle tracks reconstructed erroneously. In summary, this thesis will not only present examples of how high energy physics can greatly benefit by leveraging machine learning techniques for physics analysis and reconstruction, but will also provide guidance on how the field can prepare for the inevitable increase in machine learning applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗

Beam-ion Studies in NSTX and NSTX-U (Final Technical Report)

The confinement of high energy "fast ions" is crucial for the success of magnetic fusion as a practical energy source. The research performed by UC Irvine on NSTX and NSTX-U provided new information about this important topic in the configuration known as a "spherical tokamak." Injected neutral beams provided the fast ions (also known as "beam ions"). There were five overarching goals of the research. One objective was to measure the confinement of beam ions in order to ascertain if they were accomplishing their desired purpose of transferring their energy to the bulk plasma. A second goal was to better understand instabilities that are driven unstable by the fast ions. A related goal was to use the similarities and differences between NSTX and the DIII-D conventional tokamak to better understand the fast-ion driven instabilities. A fourth goal was to develop new instruments to measure fast ions and techniques that facilitate analysis of the data. The fifth objective was to measure and understand acceleration of beam ions by RF waves. Much was accomplished in all of these five areas. In the first, it was shown that the large magnetic moment of the beam ions did not harm confinement but instabilities known as the "sawtooth" and "long-lived mode" do. In the second area, much attention was devoted to the fast-ion instabilities known as "Alfvén eigenmodes." In particular, Alfvén eigenmode "avalanches" can cause ~ 30% of the fast ions to be lost in a single explosive burst. The comparative studies with DIII-D showed that an instability discovered on NSTX is also important in conventional tokamaks. In the fourth category, several instruments were developed, a number of effects that complicate interpretation of the data were understood, and progress toward a new method to infer the fast-ion distribution function from the data was made. In the fifth category, the measured profile of accelerated fast ions was initially much broader than theoretical predictions but subsequent improvements in the theoretical modeling achieved good agreement with the data.

43 PARTICLE ACCELERATORS↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Phase Space Reconstruction from Accelerator Beam Measurements Using Neural Networks and Differentiable Simulations

Characterizing the phase space distribution of particle beams in accelerators is a central part of accelerator understanding and performance optimization. However, conventional reconstruction-based techniques either use simplifying assumptions or require specialized diagnostics to infer high-dimensional (> $2D$) beam properties. In this Letter, we introduce a general-purpose algorithm that combines neural networks with differentiable particle tracking to efficiently reconstruct high-dimensional phase space distributions without using specialized beam diagnostics or beam manipulations. Furthermore, we demonstrate that our algorithm accurately reconstructs detailed 4D phase space distributions with corresponding confidence intervals in both simulation and experiment using a single focusing quadrupole and diagnostic screen. This technique allows for the measurement of multiple correlated phase spaces simultaneously, which will enable simplified 6D phase space distribution reconstructions in the future.

47 OTHER INSTRUMENTATION↗

Impact of the magnetic horizon on the interpretation of the Pierre Auger Observatory spectrum and composition data

The flux of ultra-high energy cosmic rays reaching Earth above the ankle energy (5 EeV) can be described as a mixture of nuclei injected by extragalactic sources with very hard spectra and a low rigidity cutoff.Extragalactic magnetic fields existing between the Earth and the closest sources can affect the observed CR spectrum by reducing the flux of low-rigidity particles reaching Earth. We perform a combined fit of the spectrum and distributions of depth of shower maximum measured with the Pierre Auger Observatory including the effect of this magnetic horizon in the propagation of UHECRs in the intergalactic space.We find that, within a specific range of the various experimental and phenomenological systematics, the magnetic horizon effect can be relevant for turbulent magnetic field strengths in the local neighbourhood in which the closest sources lieof order B$_{rms}$ ≃ (50–100) nG (20 Mpc/d$_{s}$)( 100 kpc/L$_{coh}$)$^{1/2}$, with d$_{s}$ the typical intersource separation and L$_{coh}$ the magnetic field coherence length. When this is the case,the inferred slope of the source spectrum becomes softer and can be closer to the expectations of diffusive shock acceleration, i.e., ∝ E$^{-2}$.An additional cosmic-ray population with higher source density and softer spectra, presumably also extragalactic and dominating the cosmic-ray flux at EeV energies, is also required to reproduce the overall spectrum and composition results for all energies down to 0.6 EeV.

79 ASTRONOMY AND ASTROPHYSICS↗

A mechanistic model for creep and thermal aging in Alloy 709

This report describes a physics-based model for creep and thermal aging in Alloy 709. Alloy 709 is an advanced austenitic alloy, targeted for use in future Sodium Fast Reactors (SFRs) and other advanced reactors. The material has superior high temperature properties compared to currently qualified 316 and 304 stainless steels. However, the available creep and thermal aging test database for Alloy 709 is significantly more limited compared to the historical materials. The physics-based model developed here is one way to accelerate the qualification of the material by providing more accurate long-term predictions for creep properties and thermal aging, compared to current empirical time-extrapolate techniques. The crystal plasticity finite element model is used to predict the deformation and failure of alloy 709. The same setup for the CPFE model is used in both the baseline model calibration process and the simulation campaigns for parameter inference. Specific constitutive choices are made for Alloy 709 to capture the primary deformation mechanisms. The dislocation creep formulation developed by Hu and Cocks is extended to account for coupled precipitation formation and the grain boundary cavitation model developed by Sham, Needleman, et al. is used to model grain boundary cavitation-induced failure. A novel update algorithm is proposed to render the semi-discrete constitutive update for the Sham-Needleman model unconditionally stable. A progressive calibration approach is adopted based on the observations that several types of material responses can be effectively decoupled. A surrogate model is trained based on full-fledged CPFE simulations to accelerate the forward model evaluations, and stochastic variational inference (SVI) is used to calibrate the unknown microstructural model parameters. The calibrated mechanistic model is used to predict the long-term creep life of Alloy 709, and the predictions are compared against classical empirical approaches.

36 MATERIALS SCIENCE↗

LUNA: LUT-Based Neural Architecture for Fast and Low-Cost Qubit Readout

Qubit readout is a critical operation in quantum computing systems, which maps the analog response of qubits into discrete classical states. Deep neural networks (DNNs) have recently emerged as a promising solution to improve readout accuracy . Prior hardware implementations of DNN-based readout are resource-intensive and suffer from high inference latency, limiting their practical use in low-latency decoding and quantum error correction (QEC) loops. This paper proposes LUNA, a fast and efficient superconducting qubit readout accelerator that combines low-cost integrator-based preprocessing with Look-Up Table (LUT) based neural networks for classification. The architecture uses simple integrators for dimensionality reduction with minimal hardware overhead, and employs LogicNets (DNNs synthesized into LUT logic) to drastically reduce resource usage while enabling ultra-low-latency inference. We integrate this with a differential evolution based exploration and optimization framework to identify high-quality design points. Our results show up to a 10.95x reduction in area and 30% lower latency with little to no loss in fidelity compared to the state-of-the-art. LUNA enables scalable, low-footprint, and high-speed qubit readout, supporting the development of larger and more reliable quantum computing systems.

Farooq, M. A. [Arizona State U., Tempe]↗

Machine-Learning Accelerated Studies of Materials with High Performance and Edge Computing

In the studies of materials, experimental measurements often serve as the reference to verify physics theory and modeling; while theory and modeling provide a fundamental understanding of the physics and principles behind. However, the interactions and cross validation between them have long been a challenge even to-date. Not only that inferring a physics model from experimental data is itself a difficult inverse problem, another major challenge is the orders-of-magnitude longer wall-clock time required to carry out high-fidelity computer modeling to match the timescale of experiments. We envisage that by combining high performance computing, data science, and edge computing technology, the current predicament can be alleviated, and a new paradigm of data-driven physics research will open up. For example, we can accelerate computer simulations by first performing the large-scale modeling on high performance computers and train a machine-learned surrogate model. This computationally inexpensive surrogate model can then be transferred to the computing units residing closely to the experimental facilities to perform high-fidelity simulations at a much higher throughout. The model will also be more amenable to analyzing and validating experimental observations in comparable time scales at a much lower computational cost. Further integration of these accelerated computer simulations with an outer machine learning loop can also inform and direct future experiments, while making the inverse problem of physics model inference more tractable. We will demonstrate a proof-of-concept by using a quantum Monte Carlo application, Dynamical Cluster Approximation (DCA++), to machine-learn a surrogate model and accelerate the study of quantum correlated materials.

Li, Ying Wai↗

Surrogates for Valve-Controlled Pipe Flow: Accelerating Nuclear Reactor Design

Neural surrogate models are developed to replace expensive steady-state RANS CFD simulations for valve-controlled pipe flow in nuclear reactor design. Using parametric CFD data generated with MOOSE Pronghorn across a range of valve geometry and flow conditions, three approaches are compared: a POD-based reduced-order model, a structured UNet on a cylindrical grid, and unstructured models (DeepONet and BiStride MeshGraphNet) on nondimensionalized point clouds. POD achieves the highest accuracy (99%) with fast inference but requires storing all solution snapshots, while the DeepONet and BSMS-GNN both achieve ~89% accuracy at sub-second inference, with the BSMS-GNN offering superior geometric generalizability. These surrogates enable rapid ranking of candidate valve designs and can warm-start CFD solvers to accelerate convergence, supporting agentic design iteration on the Prometheus platform.

42 - ENGINEERING↗

ATLAS Data Analysis using a Parallel Workflow on Distributed Cloud-based Services with GPUs

A new type of parallel workflow is developed for the ATLAS experiment at the Large Hadron Collider, that makes use of distributed computing combined with a cloud-based infrastructure. This has been developed for a specific type of analysis using ATLAS data, one popularly referred to as Simulation-Based Inference (SBI). The JAX library is used for the parts of the workflow to compute gradients as well as accelerate program execution using just-in-time compilation, which becomes essential in a full SBI analysis and can also offer significant speed-ups in more traditional types of analysis.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

CosmoLensNRE

Cosmology Inference from Strong Gravitational Lensing using Neural Ratio Estimation

Jarugula, Sreevani [Fermi National Accelerator Lab↗

DiHydrogen

DiHydrogen is the second version of the Hydrogen fork of the well-known distributed linear algebra library, Elemental. DiHydrogen is a GPU-accelerated distributed multilinear algebra interface with a particular emphasis on the needs of the scalable distributed deep learning training and inference. DiHydrogen is part of the Livermore Big Artificial Neural Network (LBANN) software stack.

Maruyama, Naoya↗

Interpreting and Accelerating Transformers for Jet Tagging

Attention-based transformers are ubiquitous in machine learning applications from natural language processing to computer vision. In high energy physics, one central application is to classify collimated particle showers in colliders based on the particle of origin, known as jet tagging. In this work, we study the interpretatbility and prospects for acceleration of Particle Transformer (ParT), a state-of-the-art model, leverages particle-level attention to improve jet-tagging performance. We analyzing ParT's attention maps and particle-pair correlations in the eta-phi plane, revealing intriguing features, such as a binary attention pattern that identifies critical substructure in jets. These insights enhance our understanding of the model's internal workings and learning process and hint at ways to improve its efficiency. Along these lines, we also explore low-rank attention, attention alternatives, and dynamic quantization to accelerate transformers for jet tagging. With quantization, we achieve a 50% reduction in model size and a 10% increase in inference speed without compromising accuracy. These combined efforts enhance both the performance and the interpretability of transformers in high-energy physics, opening avenues for more efficient and physics-driven model designs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Differentiable Neural Architecture, Mixed Precision and Accelerator Co-Search

Quantization, effective Neural Network architecture, and efficient accelerator hardware are three important design paradigms to maximize accuracy and efficiency. Mixed Precision Quantization is a process of assigning different precision to different Neural Network layers for optimized inference. Neural Architecture Search (NAS) is a process of automatically designing the neural network for a task and can also be extended to search for the precision of each weight and activation matrix. In this paper, we develop the following three methods: (i) Fast Differentiable Hardware-aware Mixed Precision Quantization Search method to find optimal precision, (ii) Joint Differentiable hardware-aware Architecture and Mixed Precision Quantization Co-search, (iii) Joint Accelerator, Architecture, and Precision triple co-search to find best possibilities in all the three worlds. We demonstrate the effectiveness of our proposed methods targeting Bitfusion accelerator by searching mixed precision models on MobilenetV2. We achieve better accuracy-latency trade-off models than the manually designed and previously proposed search methods.

97 MATHEMATICS AND COMPUTING↗