Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “edge inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data Center High-Temperature Liquid Cooling and Heat Reuse Techno-Economic Study: Preprint

Data centers are energy-intensive facilities with growing demands for efficiency and cost-effective operations. Smaller, more distributed edge inference data centers are expected to proliferate as AI applications require low latency closer to the user of AI tools, which presents a growing opportunity to explore the systems implications of liquid cooling on water and energy use. This study analyzes the implementation of high-temperature liquid cooling systems in a prototypical inference 1-MW data center and explores the potential for heat reuse across varying climates with a goal to optimize energy efficiency, reduce capital and operational costs, and identify opportunities for high-performance cooling and water use reduction infrastructure. This analysis evaluated configurations utilizing a peak day hourly sizing and systems performance spreadsheet to evaluate design and operational conditions from which component sizes, installed cost, operational cost, and performance metrics were determined for the Base case and the Elevated case. The techno-economic analysis included heat reuse applications across a range of heat recovery temperatures and heat rejection options. The analysis shows that high-temperature liquid cooling allows for improved energy efficiency, lower water consumption, and lower capital costs compared to traditional cooling approaches. Transitioning to elevated water inlet/outlet temperatures (50 degrees C/60 degrees C) eliminates the need for chillers, cooling towers, and heat recovery equipment in many scenarios across three distinct climate zones. This results in up to 75% capital cost savings for the cooling and heat recovery equipment, and with significantly reduced water consumption, especially in non-heat reuse applications. Heat generated from data centers can also be repurposed for space heating, domestic hot water, and other applications, and is most cost-effective when data center outlet temperatures exceed 55-60 degrees C.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Multi-Tier Autonomous Aerial Architecture for Wildfire Detection, Characterization, and Communication in Infrastructure-Denied Environments

Wildfire response depends on how fast an ignition can be confirmed and located, especially in remote regions where ground-based communication and monitoring may be limited. Geostationary sensors provide frequent observations but at kilometer-scale resolution, which is too coarse to resolve small fires in remote terrain. Ground camera networks require sightlines and infrastructure that back-country areas lack. To address these limitations, this work proposes a Multi-Tier Autonomous Wildfire Intelligence System that combines wide-area monitoring with targeted, high-resolution sensing. A solar-powered high-altitude long endurance (HALE) platform operating at approximately 60,000 ft provides persistent wide-area thermal and optical surveillance, running onboard edge inference to screen candidate ignitions and reduce false positives and downlink bandwidth. When a candidate ignition is detected, low-altitude uncrewed aircraft systems (UAS) can be deployed to conduct localized observations, including high-resolution imaging and atmospheric measurements such as wind and plume observation. By combining persistent detection with local sensing, the proposed architecture is designed to provide first responders with timely, high-resolution information about fire location and behavior to aid in emergency decision making.

Wildfire management, UAS, drones↗

Accurate model of the projected velocity distribution of galaxies in dark matter haloes

ABSTRACT We present a per cent-level accurate model of the line-of-sight velocity distribution of galaxies around dark matter haloes as a function of projected radius and halo mass. The model is developed and tested using synthetic galaxy catalogues generated with the UniverseMachine run on the Multi-Dark Planck 2 N-body simulations. The model decomposes the galaxies around a cluster into three kinematically distinct classes: orbiting, infalling, and interloping galaxies. We demonstrate that: (1) we can statistically distinguish between these three types of galaxies using only projected line-of-sight velocity information; (2) the halo edge radius inferred from the line-of-sight velocity dispersion is an excellent proxy for the three-dimensional halo edge radius; and (3) we can accurately recover the full velocity dispersion profile for each of the three populations of galaxies. Importantly, the velocity dispersion profiles of the orbiting and infalling galaxies contain five independent parameters – three distinct radial scales and two velocity dispersion amplitudes – each of which is correlated with mass. Thus, the velocity dispersion profile of galaxy clusters has inherent redundancies that allow us to perform non-trivial systematics checks from a single data set. We discuss several potential applications of our new model for detecting the edge radius and constraining cosmology and astrophysics using upcoming spectroscopic surveys.

Astronomy & Astrophysics↗

Evaluation of Vertical Patterns in Chlorophyll-A Derived From A Data Assimilating Model of Satellite-Based Ocean Color

Satellite-based sensors of ocean color have become the primary tool to infer changes in surface chlorophyll, while BGC-Argo floats are now filling the information gap at depth. Here we use BGC-Argo data to assess depth-resolved information on chlorophyll-a derived from an ocean biogeochemical model constrained by the assimilation of surface ocean color remote sensing. The data-assimilating model replicates well the general seasonality and meridional gradients in surface and depth-resolved chlorophyll-a inferred from the float array in the Southern Ocean. On average, the model tends to overestimate float-based chlorophyll, particularly at times and locations of high productivity such as the beginning of the spring bloom, subtropical deep chlorophyll maxima, and non-iron limited regions of the Southern Ocean. The highest model RMSE in the upper 50 m with respect to the float array is of 0.6 mg Chl m −3 , which should allow the detection of seasonal changes in float-based biomass (varying between 0.01 and >1 mg Chl m −3 ) but might hinder the identification of subtle changes in chlorophyll at narrow local scales. Both model and float profiling data show good agreement with in situ data from station ALOHA, with model estimates showing a slight accuracy edge in inferring depth-resolved observations. Uncertainties in float bio-optical estimates impede their use as a reliable benchmark for validation, but the general qualitative agreement between model and float data provides confidence in the ability of model to replicate biogeochemical features below the surface, where data is not directly constrained by the assimilation of satellite ocean color.

Lionel A Quintero↗

Fast Adaptive Neural Control of Resonant Extraction at Fermilab

We present progress on the development of a machine learning (ML) regulation system for third-order resonant extraction of the beam delivered to the Mu2e experiment at Fermilab. We consider classical and ML-based controllers optimized on semi-analytic simulations and provide performance comparisons for several models. Additionally, we discuss the efficiency of each model in training, which has implications for future work on adaptive control. We also discuss progress on developing optimized implementations of ML models for edge-based inference.

Whitbeck, A. [Fermilab] (ORCID:0000000342245164)↗

Fast Adaptive Neural Control of Resonant Extraction at Fermilab

We present the development of a machine learning (ML) based regulation system for third-order resonant beam extraction in the Mu2e experiment at Fermilab. Classical and ML-based controllers have been optimized using semi-analytic simulations and evaluated in terms of regulation performance and training efficiency. We compare several controller architectures and discuss the integration of neural control into an adaptive framework. We also present progress on surrogate models that predict the controller response given a spill intensity and controller action history. To enable real-time deployment, we report progress on implementing low-latency, edge-based inference suitable for hardware-constrained environments. Our results demonstrate the feasibility and advantages of ML-based control in managing complex, time-varying physical systems, with broader implications for accelerator operations and other domains requiring fast, adaptive regulation.

Berlioz, Jose Rene [Fermilab]↗

Fast Adaptive Neural Control of Resonant Extraction at Fermilab

We present progress on the development of a machine learning (ML) regulation system for third-order resonant extraction of the beam delivered to the Mu2e experiment at Fermilab. We consider classical and ML-based controllers optimized on semi-analytic simulations and provide perfor- mance comparisons for several models. Additionally, we discuss the efficiency of each model in training, which has implications for future work on adaptive control. We also discuss progress on developing optimized implementations of ML models for edge-based inference.

Whitbeck, A. [Fermilab]↗

Laser velocimetry applied to transonic and supersonic aerodynamics

Measurements obtained with laser velocimetry in a Mach 2.9 separated turbulent boundary layer and in the transonic flow past a two-dimensional airfoil section are presented and compared to data realized by conventional techniques. Agreement in mean velocities was realized where the pressure measurements could be considered reliable; however, in regions of instantaneous reverse velocities, the laser results were found to be consistent with the physics of the flow whereas the pressure data were not. Streamwise turbulence intensities are also presented. In the transonic airfoil study, velocity measurements obtained immediately outside the upper surface boundary layer of a 6-inch chord NACA 64A010 airfoil are compared to edge velocities inferred from surface pressure measurements. For free-stream Mach numbers of 0.6 and 0.8, the agreement in results was very good. "Dual scatter" optical arrangements in conjunction with a single particle, counter-type signal processor were employed in these investigations.

Johnson, D. A.↗

Laser velocimetry applied to transonic and supersonic aerodynamics

As a further demonstration of the capabilities of laser velocity in compressible aerodynamics, measurements obtained in a Mach 2.9 separated turbulent boundary layer and in the transonic flow past a two-dimensional airfoil section are presented and compared to data realized by conventional techniques. In the separated-flow study, the comparisons were made against pitot-static pressure data. Agreement in mean velocities was realized where the pressure measurements could be considered reliable; however, in regions of instantaneous reverse velocities, the laser results were found to be consistent with the physics of the flow whereas the pressure data were not. The laser data obtained in regions of extremely high turbulence suggest that velocity biasing does not occur if the particle occurrence rate is low relative to the turbulent fluctuation rate. Streamwise turbulence intensities are also presented. In the transonic airfoil study, velocity measurements obtained immediately outside the upper surface boundary layer of a 6-inch chord MACA 64A010 airfoil are compared to edge velocities inferred from surface pressure measurements. For free-stream Mach numbers of 0.6 and 0.8, the agreement in results was very good. Dual scatter optical arrangements in conjunction with a single particle, counter-type signal processor were employed in these investigations. Half-micron-diameter polystyrene spheres and naturally occurring condensed oil vapor acted as light scatterers in the two respective flows. Bragg-cell frequency shifting was utilized in the separated flow study.

Johnson, D. A.↗

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

Das, Arghya Ranjan [Purdue U.] (ORCID:000000018451↗

Neuro-Spark: A Submicrosecond Spiking Neural Networks Architecture for In-Sensor Filtering

Neuro-Spark, which is a new neuromorphic architecture with a field-programmable gate array (FPGA) implementation for ultrafast spiking neural network (SNN) inference at the edge, facilitates smart-pixel in-sensor filtering for high-energy physics experiments at the Large Hadron Collider (LHC). Utilizing the evolutionary optimization for neuromorphic systems (EONS) training method, we generate compact SNN models with 91% signal efficiency, akin to convolutional neural networks but with half the parameters. However, deploying near the detector poses a challenge because the SNN must handle a sustained input data rate exceeding 1013 GB/s. To overcome this, we propose a novel hardware architecture that uses high-level synthesis to construct a tuned architecture for the EONS-trained SNN. In addition to the analysis and validation with an AMD Xilinx Artix-A7 FPGA, our solution consumes only ç24% of FPGA LUT and flipflops. We also introduce an innovative quantization method that reduces FPGA resource utilization by ç15% without compromising accuracy. Our FPGA implementation achieves computing latency of ç10 ns for smart-pixel application inference on an edge FPGA.

Miniskar, Narasinga Rao↗

The 10 September 2025 M w 4.1 Earthquake in Northeastern Utah, United States: An Archetypal Continental Mantle Event

The 10 September 2025 M w 4.1 earthquake in northeastern Utah, United States, had a focal depth 68 km beneath sea level, which is ∼20–25 km greater than estimates of local crustal thickness, making it a rare example of a continental mantle earthquake (CME). The focal depth is well resolved from arrival-time inversion (nearest station ∼13 km away) and moment tensor inversion of regional waveforms. Similar to other CMEs in the Intermountain West, there were no obvious aftershocks or foreshocks, and the waveforms were enriched in high-frequency energy. Spectral modeling gives a stress drop of ∼80 MPa and a radiation efficiency of ∼0.08, albeit with large uncertainties. The high stress drop and low radiation efficiency are consistent with a dissipative source process such as thermal runaway. Also similar to previous Intermountain West CMEs, the event occurred along the boundary of the Archean Wyoming craton, where pressure–temperature conditions favor ductile deformation. We hypothesize that edge-driven or regional-scale mantle convection produces increased strain rates near the craton boundary that make either conventional brittle failure or thermal runaway feasible at relatively high pressure–temperature conditions. High conductivity inferred around the edge of the craton may suggest that fluids also contribute to CME occurrence.

Koper, Keith D. [Univ. of Utah, Salt Lake City, UT↗

Robust Machine Learning Inference from X-ray Absorption Near Edge Spectra through Featurization

X-ray absorption spectroscopy (XAS) is a commonly employed technique for characterizing functional materials. In particular, X-ray absorption near edge spectra (XANES) encode local coordination and electronic information, and machine learning approaches to extract this information are of significant interest. To date, most ML approaches for XANES have primarily focused on using the raw spectral intensities as input, overlooking the potential benefits of incorporating spectral transformations and dimensionality reduction techniques into ML predictions. Here, in this work, we focused on systematically comparing the impact of different featurization methods on the performance of ML models for XAS analysis. We evaluated the classification and regression capabilities of these models on computed data sets and validated their performance on previously unseen experimental data sets. Our analysis revealed an intriguing discovery: the cumulative distribution function feature achieves both high prediction accuracy and exceptional transferability. This remarkably robust performance can be attributed to its tolerance to horizontal shifts in the spectra, which is crucial when validating models using experimental data. While this work exclusively focuses on XANES analysis, we anticipate that the methodology presented here will hold promise as a versatile asset to the broader spectroscopy community.

36 MATERIALS SCIENCE↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Cognitive IoT and Edge Computing for Intrusion Detection with Federated TinyML

Internet of Things (IoT) and Edge Computing (EC) are rapidly becoming an integral part of the modern society. By 2030, there is estimated to be over 40 billion active and connected IoT devices [1]. This rapid progress also comes with a significant implication on cybersecurity. Back-end infrastructure and systems have a much broader attack than they did previously due to vulnerable IoT/EC devices being connected to wireless networks. This expanding attack surface is a growing concern because IoT/EC are increasingly being used in critical systems such as power grids, health care, and smart homes. To effectively address a problem of this scale, cognitive cyber methods—which can autonomously detect and react to cyber attacks as they develop—are needed. To address this, we bring Artificial Intelligence (AI) and Machine Learning (ML) to IoT/EC devices, using tinyML to monitor voluminous IoT data against cyber threats, and using Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. We propose a novel three-layer architecture: (1) an IoT layer for tinyML-based inference, (2) an edge layer for ML model training, and (3) a cloud layer for FL operations. Using the publicly available 11-class N-BaIoT dataset [2], we demonstrate that this architecture mitigates resource constraints at the IoT layer while improving detection accuracy over standard two-layer designs. An outlier-resistant scaler, feature reduction, and quantization enable the tinyML model to maintain detection accuracy with a reduced model size. Additionally, federated learning that only utilizes the intersection (across heterogenous devices) of the reduced feature set achieves superior detection accuracy compared to locally trained models.

Li, Mingyan [ORNL] (ORCID:0009000569532640)↗

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks

Recent years have seen an increased importance of neural network inference in edge-based scenarios, which impose size and power constraints requiring novel computing devices. These same edge scenarios may require operating over long periods of time, or exposure to extreme environments, resulting in a drift of neural network weights that cause degraded performance. In searching for ways to develop neural network approaches that perform robustly under these conditions, we propose a biologically-inspired mechanism for the dynamic adaptation of within-neuron parameters that is guided by a global context signal carrying information about perturbations and variability in incoming stimuli. Specifically, we demonstrate that adaptive voltage thresholds or neuronal time constants, when informed by a global context signal, can enable network-level mechanisms to recover from perturbed synaptic weights. Consistent with prior literature, the context-modulated approach is effective for recurrent, but not feedforward networks, by modulating network level dynamics. We demonstrate this approach successfully recovers performance in image classification tasks and spatiotemporal tracking tasks under idealized and Gaussian noise as well as for realistic perturbations from a memristive device when exposed to ionizing radiation. Finally, we discuss how this approach enables the design of robust and energy-efficient neuromorphic systems that perform well, even in resource-constrained scenarios with extreme environments such as edge processing.

context modulation↗

A 3D Implementation of Convolutional Neural Network for Fast Inference

Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.

Miniskar, Narasinga Rao↗

Inference of main ion particle transport coefficients with experimentally constrained neutral ionization during edge localized mode recovery on DIII-D

Abstract The plasma and neutral density dynamics after an edge localized mode are investigated and utilized to infer the plasma transport coefficients for the density pedestal. The Lyman-Alpha Measurement Apparatus (LLAMA) diagnostic provides sub-millisecond profile measurements of the ionization and neutral density and shows significant poloidal asymmetries in both. Exploiting the absolute calibration of the LLAMA diagnostic allows quantitative comparison to the electron and main ion density profiles determined by charge-exchange recombination, Thomson scattering and interferometry. Separation of diffusion and convection contributions to the density pedestal transport are investigated through flux gradient methods and time-dependent forward modeling with Bayesian inference by adaptation of the Aurora transport code and IMPRAD framework to main ion particle transport. Both methods suggest time-dependent transport coefficients and are consistent with an inward particle pinch on the order of 1 m s −1 and diffusion coefficient of 0.05 m 2 s −1 in the steep density gradient region of the pedestal. While it is possible to recreate the experimentally observed phenomena with no pinch in the pedestal, low diffusion in the core and high outward convection in the near scrape-off layer are required without an inward pedestal pinch.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗