Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

Applications of NASA’s LANCE Earth Observations for Ecosystem Assessments

As one of NASA's open and free data systems, NASA's Land, Atmosphere Near real-time Capability for EOS (LANCE) provides a wide range of near real-time, low latency and expedited data products from NASA and other earth science satellite missions to support users in their understanding of the value and distribution of ecosystem services. LANCE data can support the System of Environmental Economic Accounting-Ecosystem Accounting (SEEA-EA) framework by both facilitating the creation of ecosystem spatial extent and condition accounts, as well as the quantification and mapping of ecosystem services, much more quickly than routine data processing allows. During the past 13 years, LANCE near real-time satellite data products (e.g., surface reflectance, albedo, vegetation height, thermal anomalies, soil moisture, snow cover etc.) have been used to produce ecosystem-related indicators such as vegetation indexes, biomass, land cover maps, fire, and flood products. These ecosystem-related indicators can be integrated into ecosystem services models, tools, and platforms to monitor land cover and land use changes in ecosystem composition in support of land and natural resource management. When using ecosystem services accounts in decision-making, low latency is essential because having up-to-date information can be critical to support decisions being made. With adequate spatial resolution, long time series length, and low latency data products, LANCE would promote the sustainable use of land and natural resources.

Tian Yao↗

Why Near Real-time/Low Latency is Important to Monitor the Changing World

An essential factor for remote sensing data products in impacting decision making is latency, or the time between earth observation and data are available to users. In many applications areas, latency plays an important or even decisive role where low latency earth observations help people to make timely, data-based decisions. Within the open and free NASA resources, NASA’s Land, Atmosphere Near real-time Capability for Earth Observing System (EOS) (LANCE) supports users interested in monitoring a wide variety of natural and man-made phenomena using near real-time (NRT) data products that are made available much quicker than routine processing allows. The combination of all available LANCE satellite products provides global coverage at multiple times per day, which makes it possible to meet user needs in various areas of applications including water resources, agriculture, air quality, wildland fire and many other disasters monitoring and management. As one of the prime users of LANCE, NASA’s Earth Science Applied Sciences Program promotes the use of LANCE data products to demonstrate applications in decision making and facilitates the end-user feedback to the science team to improve data products. One of the most critical applications that we expect data to be processed as close to the user as possible is wildland file response and management. User feedback indicates that LANCE NRT fire data products within 3 hours latency would meet the needs of the wildland fire community. Other examples of earth science application areas for which low latency is particularly important include detecting volcanic eruptions, early warning of disasters, tracking extreme weather events, and monitoring air quality.

Tian Yao↗

Reconfigurable Network Slicing Orchestration in Network Function Virtualization Compatible Operational Technology Environment

The ongoing transition to Industry 4.0, which is characterized by increased inter-connectivity of cyber-physical systems, requires having time-sensitive, high throughput, and secure transfer of critical data in industrial sites. In this context, network slicing emerges as a critical tool to ensure timely data delivery by provisioning the network resources to cater to specific applications’ requirements and mitigating potential cyber attacks. To address these challenges, this paper aims to tackle two key questions essential for the successful implementation of network slicing in industrial environments. First, it investigates architectural considerations for developing a network infrastructure capable of supporting network slicing functionalities effectively. The proposed approach significantly improves deployment efficiency over traditional manual configurations. Second, it delves into the automated orchestration process, elucidating the steps and components involved in transitioning from a static network management approach to dynamically leverage network function virtualization schemes for creating network slices in ad-hoc manner. The system demonstrates high throughput suitable for production-level solutions and maintains exceptionally low latency, making it ideal for ultra-reliable low-latency communications. Even with increased network demands, the system remains stable, with effective Quality of Service (QoS) management, ensuring reliable performance under varying conditions. The proposed architecture outlines the necessary components, services, and communication protocols required for a production-level orchestrator for network segmentation in SCADA environments.

Rodiles Delgado, Brian G.↗

A Status Update for the FLASHFlux Working Group

This presentation provides an overview of the progress made by the FLASHFlux working group within the CERES Science Team. FLASHFlux is maintaining operations, continuing validation, migration production code for future production systems, and evaluating new inputs for their impacts on the current radiative flux data products. FLASHFlux data products continue to be served to the community and comprise the low latency solar and thermal infrared data products provided to the energy and agricultural communities at relatively low latency for global gridded data fluxes. (< 7days).

surface radiative flux↗

STRIPE: Remote Driving Using Limited Image Data

Driving a vehicle, either directly or remotely, is an inherently visual task. When heavy fog limits visibility, we reduce our car's speed to a slow crawl, even along very familiar roads. In teleoperation systems, an operator's view is limited to images provided by one or more cameras mounted on the remote vehicle. Traditional methods of vehicle teleoperation require that a real time stream of images is transmitted from the vehicle camera to the operator control station, and the operator steers the vehicle accordingly. For this type of teleoperation, the transmission link between the vehicle and operator workstation must be very high bandwidth (because of the high volume of images required) and very low latency (because delayed images can cause operators to steer incorrectly). In many situations, such a high-bandwidth, low-latency communication link is unavailable or even technically impossible to provide. Supervised TeleRobotics using Incremental Polyhedral Earth geometry, or STRIPE, is a teleoperation system for a robot vehicle that allows a human operator to accurately control the remote vehicle across very low bandwidth communication links, and communication links with large delays. In STRIPE, a single image from a camera mounted on the vehicle is transmitted to the operator workstation. The operator uses a mouse to pick a series of 'waypoints' in the image that define a path that the vehicle should follow. These 2D waypoints are then transmitted back to the vehicle, where they are used to compute the appropriate steering commands while the next image is being transmitted. STRIPE requires no advance knowledge of the terrain to be traversed, and can be used by novice operators with only minimal training. STRIPE is a unique combination of computer and human control. The computer must determine the 3D world path designated by the 2D waypoints and then accurately control the vehicle over rugged terrain. The human issues involve accurate path selection, and the prevention of disorientation, a common problem across all types of teleoperation systems. STRIPE is the only semi-autonomous teleoperation system that can accurately follow paths designated in monocular images on varying terrain. The thesis describes the STRIPE algorithm for tracking points using the incremental geometry model, insight into the design and redesign of the interface, an analysis of the effects of potential errors, details of the user studies, and hints on how to improve both the algorithm and interface for future designs.

REMOTE CONTROL↗

Advancing Solar Energetic Particle Forecasting

With growing interest from the aviation and satellite industries, and for NASA’s upcoming Artemis lunar missions, the need for improved scientific understanding and accurate forecasting of solar energetic particle events has never been stronger. In this paper we discuss the observational, validation and model transition support required to achieve these goals. Well-calibrated, high-quality energetic electron, proton, and ion measurements are essential. Expansions to the fields of view offered by current X-ray, extreme ultraviolet and coronagraph instruments, to obtain increased coverage of the solar corona and heliosphere, from vantage points off the Sun-Earth line, are desired for model input. New observations of suprathermal particles are needed to characterize seed particle distributions and low latency space-based observations of solar radio emissions are also desired. Together, this observational suite should offer high cadence, low latency, reliable and accurate space weather data streams. Consistent, extensive and quantitative model validation is required to assess scientific advancements and pave the way for models transitioning to real-time forecast operations. Model performance and skill should be compared to observations and to current operational forecasting baselines. Finally, resources are required to support the significant effort of transitioning mature models into forecast operations.

solar energetic particles↗

Toward Wireless Smart Grid Communications: An Evaluation of Protocol Latencies in an Open-Source 5G Testbed

Fifth-generation networks promise wide availability of wireless communication with inherent security features. The 5G standards also outline access for different applications requiring low latency, machine-to-machine communication, or mobile broadband. These networks can be advantageous to numerous applications that require widespread and diverse communications. One such application is found in smart grids. Smart grid networks, and Operational Technology (OT) networks in general, utilize a variety of communication protocols for low-latency control, data monitoring, and reporting at every level. Transitioning these network communications from wired Wide Area Networks (WANs) to wireless communication through 5G can provide additional benefits to their security and network configurability. However, introducing these wireless capabilities may also result in a degradation of network latency. In this paper, we propose utilizing 5G for smart grid communications, and we evaluate the latency impacts of encapsulating GOOSE, Modbus, and DNP3 for transmission over a 5G network. The OpenAirInterface open-source library is utilized to deploy an in-lab 5G Core Network and gNB for testing with off-the-shelf User Equipment (UE). This creates an effective 5G test platform for experimenting with different OT protocols such as GOOSE. The results are validated by measuring two different Intelligent Electronic Devices’ contact closure times for each network configuration. These tests are also conducted for varying packet sizes in order to isolate different sources of network latency. Our study outlines the latency impact of communication over 5G for time-critical and non-critical applications regarding their transition toward private 5G-based OT network implementations. The conducted experiments illustrate that in the case of GOOSE packets, simple encapsulation may exceed the protocol’s time-critical nature, and, therefore, additional measures must be taken to ensure a viable transition of GOOSE to 5G services. However, non-critical applications are shown to be viable for migration to 5G.

42 ENGINEERING↗

Electron cyclotron emission detection of neoclassical tearing modes for control for ITER

Successful operation of ITER requires control of magnetic instabilities including neoclassical tearing modes (NTMs) that can degrade confinement and lead to disruption. Low latency detection by electron cyclotron emission (ECE) diagnostics has been demonstrated in a few current experiments. Using a synthetic diagnostic, we demonstrate low latency NTM detection for ITER with plasmas described by ITER IMAS database scenarios and with realistic limitations imposed on the instrumentation by these high temperature scenarios. 2/1 NTMs are detected 430 ms after magnetic island seeding and before island locking. The radiometer configuration was optimized using simulation, and the smallest detectable island size was explored. Island sizes of ∼3 cm are detectable at the 2/1 surface. The simulated signals incorporate recent physics models for island growth and rotation, which show early locking and continued island growth after locking and before disruption. This work determines limits for ITER ECE spatial resolution imposed by relativistic broadening of channels, which informs hardware design. Real-time detection is demonstrated in hardware that is required by ITER, including on an NI PXI-7853R FPGA system. Development of a synthetic diagnostic and details of the hardware will be discussed.

Cyclotron radiation↗

A 3D Implementation of Convolutional Neural Network for Fast Inference

Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.

Miniskar, Narasinga Rao↗

DENOVA: Deduplication Extended NOVA File System

This paper shows mathematically and experimentally that inline deduplication is not suitable for file systems on ultra-low latency Intel Optane DC PM devices in terms of performance, and proposes DeNova, an offline deduplication specially designed for log-structured NVM file systems such as NOVA. DeNova offers high-performance and low-latency I/O processing and executes deduplication in the background without interfering with foreground I/Os. DeNova employs DRAM-free persistent deduplication metadata, favoring CPU cache line, and ensures failure consistency on any system failure. We implement DeNova in the NOVA file system. Evaluation with DeNova confirms a negligible performance drop of baseline NOVA of less than 1%, while gaining high storage space savings. Extensive experiments show DeNova is failure consistent in all failure scenario cases.

Khan, Awais↗

ML–Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement Regimes

Accurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. To overcome this bottleneck, we present an in-situ machine learning (ML) inference engine implemented on an FPGA for real-time quantum state discrimination. Our design performs inference directly on digitized readout signals with 40 ns latency, supports both qubit and qutrit readout, and enables conditional operations without host-side intervention. This capability is critical for MCM and for feedback-driven protocols such as quantum error correction. We validate the system on superconducting transmon hardware, demonstrating robust discrimination fidelity across multiple qubit and qutrit channels. We further demonstrate conditional qutrit logic driven by FPGA-resident classification, highlighting the potential of low-latency ML-on-FPGA control for NISQ applications and scalable fault-tolerant quantum computing.

Vora, Neel [Lawrence Berkeley National Laboratory ↗

SatCORPS Global Cloud Composite (GCC): the Design and Delivery of A High Quality, High Resolution, Global Cloud Product Available in Near-Real Time

The NASA Satellite ClOud and Radiation Property retrieval System (SatCORPS) supports the development of an analysis ready and cloud-optimized data transformation pipeline and geospatial service enablement of a global cloud composite (GCC) product derived from global geostationary satellite imagery. This geospatial service will be available at high temporal and spatial resolution via the SatCORPS web mapping application for visualization and analysis as well as direct ingestion to common geospatial software and custom programming. The resulting global cloud composite products from the processing pipeline can then be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. Near real time global observations are created through the composition of five geostationary satellites that provides modelling and forecasting communities with the capability to provide high quality and timely information to start the projection process. The Global Cloud Composite product combines information from geostationary satellites, GOES-16, GOES-17, Himawari-8, Meteosat-11 and Meteosat-9 to create a single global composite netcdf file and images using the different products within the netcdf file. The SatCORPS team, though our Global Cloud Composite (GCC) product and web-based visualization tools including Geographical Information System (GIS) services provide near real time global cloud product information to both automated processes and traditional web users that is timely and high quality derived from geostationary satellites. The Global Cloud Composite product takes advantage of the scalable processing resources provided by the AWS batch service to provide new composites every thirty minutes. Because information from each of the low earth orbiting satellites is available on schedules tuned to the specific satellite, the processing algorithm temporally composites the final dataset as each satellite’s information becomes available. The SatCORPS team has leveraged our experience using Amazon Web Services (AWS) to build a low latency high availability tool that allows end users both human and automated to acquire high quality and high-resolution Geostationary Earth Orbiting (GEO) information at zero cost to the end user. This presentation will describe how we architected and implemented the service as well as lessons learned based on our experiences both developing and operating the system. The lessons learned include how we integrated multiple services including Amazon Batch, Amazon S3 and Amazon Lambda service to create a low cost but high-performance processing system that is capable of identifying and processing the most appropriate satellite overpass information into global cloud composites. We will also describe our web-based tools including our Geographic Information System that can be used for visualization and analysis. The products from the processing can be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. The SatCORPS Global Composite Cloud product provides sophisticated global composited cloud research products with very low latency that we see that as filling a rapidly growing need in the research and modelling community with no up-front nor ongoing costs associated with downloading or using the information.

AWS AMCE SMCE GCC SATCORPS GLOBAL CLOUD COMPOSITE ↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Control coordination in inverter‐based microgrids using AoI‐based 5G schedulers

Abstract A coordinated set point automatic adjustment with correction enabled (C‐SPAACE) framework that uses 5G communication for real‐time control coordination between inverter‐based resources (IBR) in microgrids is proposed. Utilising slicing capability, 5G offers low‐latency communication to C‐SPAACE under normal conditions. However, given the multitude of power grid use cases, a certain 5G slice for C‐SPAACE may have access only to limited radio spectrum resources, which if not managed well, greatly undermines the communication needs of C‐SPAACE framework. Thus, optimally scheduling the available spectrum resources among IBRs in a sliced 5G network‐based C‐SPAACE framework becomes a critical problem. To address this issue, the authors utilise a novel age of information (AoI) metric and designs an AoI‐based 5G scheduler to provide low‐latency communication to C‐SPAACE. Following this, a co‐simulation environment is designed using PSCAD/EMTDC and Python to simulate a microgrid supported by 5G communication. Time‐domain simulation case studies are performed using the proposed co‐simulation environment to evaluate the performance of C‐SPAACE using 5G with both AoI‐based and other baseline (non‐AoI) schedulers.

Choudhury, Biplav↗

An open, parallel I/O computer as the platform for high-performance, high-capacity mass storage systems

APTEC Computer Systems is a Portland, Oregon based manufacturer of I/O computers. APTEC's work in the context of high density storage media is on programs requiring real-time data capture with low latency processing and storage requirements. An example of APTEC's work in this area is the Loral/Space Telescope-Data Archival and Distribution System. This is an existing Loral AeroSys designed system, which utilizes an APTEC I/O computer. The key attributes of a system architecture that is suitable for this environment are as follows: (1) data acquisition alternatives; (2) a wide range of supported mass storage devices; (3) data processing options; (4) data availability through standard network connections; and (5) an overall system architecture (hardware and software designed for high bandwidth and low latency). APTEC's approach is outlined in this document.

Abineri, Adrian↗

Network Slicing for Federated Learning in Operational Technology Environment

Industrial Control Systems (ICS) and Supervisory Control and Data Acquisition (SCADA) environments are essential to modern infrastructure, facing challenges in ensuring low-latency, high-throughput communication while mitigating cyber threats. This paper presents a framework integrating Federated Learning (FL) and network slicing with Quality of Service (QoS) to enable real-time monitoring without disrupting OT operations. Leveraging digital twin technology and Network Function Virtualization (NFV), the architecture supports predictive analytics and Industry 4.0 requirements. FL facilitates decentralized model training, preserving data privacy and scalability, though it introduces potential throughput constraints. Network slicing addresses this by creating dedicated virtualized segments optimized for performance and security. Advanced fault tolerance at the container and instance levels enhances system reliability. The proposed architecture ensures high throughput, low latency, and secure orchestration for real-time anomaly detection in OT networks. Performance evaluations validate its efficiency in throughput, deployment, and learning accuracy, providing a robust foundation for future ICS automation and data-driven decision-making.

Delgado, Brian G. Rodiles [University of Texas at ↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗