Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “edge inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Voltage Inference Framework for Real-Time Observability in Active Distribution Grids

Active distribution grids are gaining traction to meet the growing environmental, socio-economic, and sustainability targets. Various advanced smart grid technologies facilitate the integration of Distributed Energy Resources (DERs) by supporting the bi-directional power flow. The limited observability of distribution grids, primarily related to their location at the very edge of power system infrastructure, brings challenges to optimal grid management. Moreover, only a limited number of measurements at regular intervals are usually available. This paper presents a novel inference framework, referred to as “Voltage Inference”, to overcome the observability issues. The proposed framework employs a prediction step based on the Multivariate Taylor series approximation, followed by a corrector step that minimizes the estimation error to infer the otherwise unknown voltages from the available measurements. Furthermore, numerical results on the IEEE 13-bus test feeder validate the accuracy and computational performance of the proposed framework.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

Real-Time Inference For MI/RR Deblending

The Fermilab Main Injector (MI) and Recycler Ring (RR) share a common beam loss monitor (BLM) system, making loss events difficult to attribute to their source machine when beam is present in both simultaneously. The Real-time Edge AI for Distributed Systems (READS) project addresses this by deblending BLM readings in real time using machine learning (ML). The current FPGA based implementation meets the sub-3 ms latency requirement but carries a resource intensive hls4ml development cycle, motivating exploration of GPU based deployment. This paper characterizes inference latency on an NVIDIA Jetson Orin Nano and introduces a packet organization scheme for assembling synchronized event frames from seven distributed BLM DAQ streams. Using a Python based DAQ simulation with injected timing jitter in place of unavailable live beam data, the pipeline achieved an average end to end latency of 0.456 ms (σ = 0.122 ms) across 167,000 test frames, comfortably meeting the timing constraint. Early outliers were attributed to TensorRT warm-up rather than steady state limitations, suggesting GPU based inference is a viable alternative to the existing FPGA implementation.

Yu, Kellen [Cornell U.]↗

A Machine Learning Framework to Predict Images of Edge-on Protoplanetary Disks

The physical structure and properties of protoplanetary disks are typically derived from spatially resolved disk images. Edge-on disks in particular provide an important view point on the vertical structure and degree of settling of disks. Such analyses rely on radiative transfer (RT) calculations that are generally computationally intensive due to the high optical depth of disks. Here we present a machine learning framework that has the potential to dramatically speed up the forward modeling process by approximating the results of RT calculations. This framework, trained on an initial set of RT calculations, utilizes an autoencoder neural network to enable the generation of synthetic scattered light images of edge-on disks directly from a set of physical parameters. We demonstrate that this framework generates synthetic images 2–3 orders of magnitude faster than using RT calculations. These machine learning-generated images appear to approximate the RT images well, in particular preserving their size and shape. We also find a strong correlation between the latent space representations of the generated disk images and several of their associated physical parameters. Finally, we discuss potential changes to the framework, such as methods to further improve the image quality, extending the framework to multiple wavelengths, and inverting the process to infer physical parameters from observed images. Overall, these new tools have the potential to enable a more efficient and uniform analysis of edge-on disk properties and the initial conditions of planet formation.

79 ASTRONOMY AND ASTROPHYSICS↗

DriveSense: A Noise-Resilient Framework for Driving Mode Identification

Accurate drive mode classification is essential for enhancing the reliability and predictive maintenance of heavy-duty electric trucks. This study proposes a novel fuzzy logic-based framework, DriveSense, for real-time drive mode classification, addressing key challenges such as sensor noise, transitional behaviors, and computational efficiency. The proposed approach integrates a two-stage filtering pipeline, combining adaptive outlier removal and a dynamic Kalman filter to enhance data quality. A fuzzy inference system with smoothened trapezoidal membership functions is then applied to classify driving modes into standstill, constant speed, acceleration, and deceleration while mitigating the effects of noise and edge cases. Performance evaluation using real-world and simulated drive cycles demonstrates significant improvements in classification accuracy (up to 97.8%), F1-score (up to 0.97), and robustness against noise, while reducing false positives. Comparative analysis against baseline models, demonstrates DriveSense’s superior accuracy and generalizability across diverse driving patterns. The framework’s lightweight and interpretable fuzzy inference engine operates with low computational latency, ensuring compatibility with real-time embedded systems typical of heavy-duty electric trucks. Moreover, DriveSense models transitional behaviors through overlapping fuzzy sets and adaptive borderline classification logic, enabling smooth identification of subtle shifts such as rolling stops or gradual deceleration. These results highlight DriveSense’s potential to enhance predictive maintenance strategies, reduce downtime, and support scalable, fleet-wide diagnostics.

Kumar, Praveen [Oak Ridge National Laboratory (ORN↗

Neutron backscatter edges as a diagnostic of burn propagation

High gain in hotspot-ignition inertial confinement fusion (ICF) implosions requires the propagation of thermonuclear burn from a central hotspot to the surrounding cold dense fuel. As ICF experiments enter the burning plasma regime, diagnostic signatures of burn propagation must be identified. In previous work [A. J. Crilly et al., Phys. Plasmas 27(1), 012701 (2020)], it has been shown that the spectral shape of the neutron backscatter edges is sensitive to the dense fuel hydrodynamic conditions. The backscatter edges are prominent features in the ICF neutron spectrum produced by the 180° scattering of primary deuterium–tritium fusion neutrons from ions. In this work, synthetic neutron spectra from radiation-hydrodynamics simulations of burning ICF implosions are used to assess the backscatter edge analysis in a propagating burn regime. Significant changes to the edge's spectral shape are observed as the degree of burn increases, and a simplified analysis is developed to infer scatter-averaged fluid velocity and temperature. The backscatter analysis offers direct measurement of the increased dense fuel temperatures that result from burn propagation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Upgrade of the Lyman-alpha diagnostic system on DIII-D for main chamber edge neutral studies

The LLAMA (Lyman Alpha Measurement Apparatus) pinhole camera diagnostic had previously been deployed on DIII-D to measure radial profiles of the Lyman-α (Ly-α) deuterium neutral line brightness across the plasma boundary in the lower chamber to infer neutral deuterium density and ionization rate profiles. This system has recently been upgraded with a new diagnostic head, named ALPACA, that also encloses two pinhole cameras and duplicates the LLAMA views in the upper chamber. Similar to LLAMA, ALPACA provides two times 20 lines of sight, viewing the plasma edge on the inboard and outboard sides with a radial resolution of ~2.5 cm (FWHM) and an effective time resolution of ~1 ms that allows for the investigation of inter-ELM dynamics. The extended Ly-α system provides better coverage to study neutrals in experiments with various plasma shapes utilizing both the upper and lower divertors. Furthermore, post-campaign calibration of the LLAMA diagnostic has successfully been demonstrated for the first time. This was facilitated by various upgrades to the calibration set-up and detailed measurements of the emissivity distribution of the Ly-α calibration source using a pinhole collimator. It was found that the sensitivity of the inboard LLAMA pinhole camera was reduced by a factor of 2.0 ± 0.2 over the course of six months of plasma operation in 2021. In conclusion, the upgraded Ly-α system, equipped with improved absolute calibration, will provide key input for neutral fueling and pedestal particle transport studies and for 2D edge transport code validation on the DIII-D tokamak.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exploration of Real Time Inference for MI-RR Deblending on GPU/TPU Systems

The Fermilab Main Injector (MI) and Recycler Ring (RR) share a common beam loss monitor (BLM) system, making loss events difficult to attribute to their source machine when beam is present in both simultaneously. The Real-time Edge AI for Distributed Systems (READS) project addresses this by deblending BLM readings in real time using machine learning (ML). The current FPGA based implementation meets the sub-3 ms latency requirement but carries a resource intensive hls4ml development cycle, motivating exploration of GPU based deployment. This paper characterizes inference latency on an NVIDIA Jetson Orin Nano and introduces a packet organization scheme for assembling synchronized event frames from seven distributed BLM DAQ streams. Using a Python based DAQ simulation with injected timing jitter in place of unavailable live beam data, the pipeline achieved an average end to end latency of 0.456 ms (σ = 0.122 ms) across 167,000 test frames, comfortably meeting the timing constraint. Early outliers were attributed to TensorRT warm-up rather than steady state limitations, suggesting GPU based inference is a viable alternative to the existing FPGA implementation.

Yu, Kellen [Fermilab; Cornell U.]↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗

Seismic Waveform Inversion Capability on Resource-Constrained Edge Devices

Seismic full wave inversion (FWI) is a widely used non-linear seismic imaging method used to reconstruct subsurface velocity images, however it is time consuming, has high computational cost and depend heavily on human interaction. Recently, deep learning has accelerated it’s use in several data-driven techniques, however most deep learning techniques suffer from overfitting and stability issues. In this work, we propose an edge computing-based data-driven inversion technique based on supervised deep convolutional neural network to accurately reconstruct the subsurface velocities. Deep learning based data-driven technique depends mostly on bulk data training. In this work, we train our deep convolutional neural network (DCN) (UNet and InversionNet) on the raw seismic data and their corresponding velocity models during the training phase to learn the non-linear mapping between the seismic data and velocity models. The trained network is then used to estimate the velocity models from new input seismic data during the prediction phase. The prediction phase is performed on a resource-constrained edge device such as Raspberry Pi. Raspberry Pi provides real-time and on-device computational power to execute the inference process. In addition, we demonstrate robustness of our models to perform inversion in the presence on noise by performing both noise-aware and no-noise training and feeding the resulting trained models with noise at different signal-to-noise (SNR) ratio values. We make great efforts to achieve very feasible inference times on the Raspberry Pi for both models. Specifically, the inference times per prediction for UNet and InversionNet models on Raspberry Pi were 22 and 4 s respectively whilst inference times for both models on the GPU were 2 and 18 s which are very comparable. Finally, we have designed a user-friendly interactive graphical user interface (GUI) to automate the model execution and inversion process on the Raspberry Pi.

Manu, Daniel (ORCID:0000000154982677)↗

Machine-Learning Accelerated Studies of Materials with High Performance and Edge Computing

In the studies of materials, experimental measurements often serve as the reference to verify physics theory and modeling; while theory and modeling provide a fundamental understanding of the physics and principles behind. However, the interactions and cross validation between them have long been a challenge even to-date. Not only that inferring a physics model from experimental data is itself a difficult inverse problem, another major challenge is the orders-of-magnitude longer wall-clock time required to carry out high-fidelity computer modeling to match the timescale of experiments. We envisage that by combining high performance computing, data science, and edge computing technology, the current predicament can be alleviated, and a new paradigm of data-driven physics research will open up. For example, we can accelerate computer simulations by first performing the large-scale modeling on high performance computers and train a machine-learned surrogate model. This computationally inexpensive surrogate model can then be transferred to the computing units residing closely to the experimental facilities to perform high-fidelity simulations at a much higher throughout. The model will also be more amenable to analyzing and validating experimental observations in comparable time scales at a much lower computational cost. Further integration of these accelerated computer simulations with an outer machine learning loop can also inform and direct future experiments, while making the inverse problem of physics model inference more tractable. We will demonstrate a proof-of-concept by using a quantum Monte Carlo application, Dynamical Cluster Approximation (DCA++), to machine-learn a surrogate model and accelerate the study of quantum correlated materials.

Li, Ying Wai↗

Precision Measurement of the Microwave Dielectric Loss of Sapphire in the Quantum Regime with Parts-per-Billion Sensitivity

Dielectric loss is known to limit state-of-the-art superconducting qubit lifetimes. Recent experiments imply upper bounds on bulk dielectric loss tangents on the order of 100 parts per billion but because these inferences are drawn from fully fabricated devices with many loss channels, these experiments do not definitely implicate or exonerate the dielectric. To resolve this ambiguity, we devise a measurement method capable of separating and resolving bulk dielectric loss with a sensitivity at the level of 5 ×10 –9 . The method, which we call the dielectric dipper, involves the in situ insertion of a dielectric sample into a high-quality microwave cavity mode. Smoothly varying the participation of the sample in the cavity mode enables a differential measurement of the dielectric loss tangent of the sample. The dielectric dipper can probe the low-power behavior of dielectrics at cryogenic temperatures and does so without the need for any lithographic process, enabling controlled comparisons of substrate materials and processing techniques. We demonstrate the method with measurements of sapphire grown by edge-defined film-fed growth (EFG) in comparison to high-grade sapphire grown by the heat-exchanger method (HEMEX). For EFG sapphire, we infer a bulk loss tangent of 63⁢(8) ×10 –9 and a substrate-air interface loss tangent of 15⁢(3) ×10 –4 (assuming a sample surface thickness of 3 nm). For a typical transmon, this bulk loss tangent would limit device quality factors to Q ≲20 ×10 6 , suggesting that bulk loss is likely the dominant loss mechanism in the longest-lived transmons on sapphire. We also demonstrate this method on HEMEX sapphire and bound its bulk loss tangent to be less than 19⁢(6) ×10 –9 . As this bound is about 3 times smaller than the bulk loss tangent of EFG sapphire, the use of HEMEX sapphire as a substrate would lift the bulk dielectric coherence limit of a typical transmon qubit to several milliseconds.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Conceptual study on using Doppler backscattering to measure magnetic pitch angle in tokamak plasmas

We introduce a new approach to measure the magnetic pitch angle profile in tokamak plasmas with Doppler backscattering (DBS), a technique traditionally used for measuring flows and density fluctuations. The DBS signal is maximised when its probe beam's wavevector is perpendicular to the magnetic field at the cutoff location, independent of the density fluctuations [Hillesheim \emph{et al} 2022 \emph{Nucl. Fusion} \textbf{55} 073024]. Hence, if one could isolate this effect, DBS would then yield information about the magnetic pitch angle. By varying the toroidal launch angle, the DBS beam reaches cutoff with different angles with respect to the magnetic field, but with other properties remaining similar. Hence, the toroidal launch angle which gives maximum backscattered power is thus that which is matched to the pitch angle at the cutoff location, enabling inference of the magnetic pitch angle. We performed systematic scans of the DBS toroidal launch angle for repeated DIII-D tokamak discharges. Experimental DBS data from this scan were analysed and combined with Gaussian beam-tracing simulations using the Scotty code [Hall-Chen \emph{et al} 2022 \emph{Plasma Phys. Control. Fusion} \textbf{64} 095002]. The pitch-angle inferred from DBS is consistent with that from magnetics-only and motional-Stark-effect-constrained (MSE) equilibrium reconstruction in the edge. In the core, the pitch angles from DBS and magnetics-only reconstructions differ by one to two degrees, while simultaneous MSE measurements were not available. The uncertainty in these measurements was under a degree; we show that this uncertainty is primarily due to the error in toroidal steering, the number of toroidally separated measurements, and shot-to-shot repeatability. We find that the error of pitch-angle measurements can be reduced by optimising the poloidal launch angle and initial beam properties. Since DBS has high spatial and temporal resolutions, is non-perturbative, does not require neutral beams, and is likely robust to neutron damage of and debris on the first mirrors, using DBS to measure the pitch angle in future fusion energy systems is especially appealing.

beam tracing↗

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML) and silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

Gonski, Julia [SLAC]↗

Impurity transport studies at the HSX stellarator using active and passive CVI spectroscopy

The transport of carbon impurities has been studied in the helically symmetric stellarator experiment (HSX) using active and passive charge exchange recombination spectroscopy (CHERS). For the analysis of the CHERS signals, the STRAHL impurity transport code has been re-written in the python programming language and optimized for the application in stellarators. In addition, neutral hydrogen densities both along the NBI line of sight as well as for the background plasma have been calculated using the FIDASIM code. By using the basinhopping algorithm to minimize the difference between experimental and predicted active and passive signals, significant levels of impurity diffusion are observed. In this work, comparisons with neoclassical calculations from DKES/PENTA show that the inferred levels exceed the neoclassical transport by about a factor of four in the core and more than 100 times towards the plasma edge, thus indicating a high level of anomalous transport. This observation is in agreement with experimental heat diffusivites determined from a power balance analysis which exhibits strong anomalous transport as well.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

‘Flux+Mutability’: a conditional generative approach to one-class classification and anomaly detection

Abstract Anomaly Detection is becoming increasingly popular within the experimental physics community. At experiments such as the Large Hadron Collider, anomaly detection is growing in interest for finding new physics beyond the Standard Model. This paper details the implementation of a novel Machine Learning architecture, called Flux+Mutability, which combines cutting-edge conditional generative models with clustering algorithms. In the ‘flux’ stage we learn the distribution of a reference class. The ‘mutability’ stage at inference addresses if data significantly deviates from the reference class. We demonstrate the validity of our approach and its connection to multiple problems spanning from one-class classification to anomaly detection. In particular, we apply our method to the isolation of neutral showers in an electromagnetic calorimeter and show its performance in detecting anomalous dijets events from standard QCD background. This approach limits assumptions on the reference sample and remains agnostic to the complementary class of objects of a given problem. We describe the possibility of dynamically generating a reference population and defining selection criteria via quantile cuts. Remarkably this flexible architecture can be deployed for a wide range of problems, and applications like multi-class classification or data quality control are left for further exploration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Towards Compact Neural Networks via End-to-End Training: A Bayesian Tensor Approach with Automatic Rank Determination

Post-training model compression can reduce the inference costs of deep neural networks, but uncompressed training still consumes enormous hardware resources and energy. To enable low-energy training on edge devices, it is highly desirable to directly train a compact neural network from scratch with a low memory cost. Low-rank tensor decomposition is an effective approach to reduce the memory and computing costs of large neural networks. However, directly training low-rank tensorized neural networks is a very challenging task because it is hard to determine a proper tensor rank a priori, and the tensor rank controls both model complexity and accuracy. Here, this paper presents a novel end-to-end framework for low-rank tensorized training. We first develop a Bayesian model that supports various low-rank tensor formats (e.g., CANDECOMP/PARAFAC, Tucker, tensor-train, and tensor-train matrix) and reduces neural network parameters with automatic rank determination during training. Then we develop a customized Bayesian solver to train large-scale tensorized neural networks. Our training methods shows orders-of-magnitude parameter reduction and little accuracy loss (or even better accuracy) in the experiments. On a very large deep learning recommendation system with over 4.2 ×10 9 model parameters, our method can reduce the parameter number to 1.6 ×10 5 automatically in the training process (i.e., by 2.6 ×10 4 times) while achieving almost the same accuracy. Code is available at https://github.com/colehawkins/bayesian-tensor-rank-determination.

compact neural networks↗

Real-Time Edge AI for Distributed Systems (READS): Progress on Beam Loss De-Blending for the Fermilab Main Injector and Recycler

The Fermilab Main Injector enclosure houses two accelerators, the Main Injector and Recycler. During normal operation, high intensity proton beams exist simultaneously in both. The two accelerators share the same beam loss monitors (BLM) and monitoring system. Beam losses in the Main Injector enclosure are monitored for tuning the accelerators and machine protection. Losses are currently attributed to a specific machine based on timing. However, this method alone is insufficient and often inaccurate, resulting in more difficult machine tuning and unnecessary machine downtime. Machine experts can often distinguish the correct source of beam loss. This suggests a machine learning (ML) model may be producible to help de-blend losses between machines. Work is underway as part of the Fermilab Real-time Edge AI for Distributed Systems Project (READS) to develop a ML empowered system that collects streamed BLM data and additional machine readings to infer in real-time, which machine generated beam loss.

43 PARTICLE ACCELERATORS↗