Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Network Slicing for Federated Learning in Operational Technology Environment

Industrial Control Systems (ICS) and Supervisory Control and Data Acquisition (SCADA) environments are essential to modern infrastructure, facing challenges in ensuring low-latency, high-throughput communication while mitigating cyber threats. This paper presents a framework integrating Federated Learning (FL) and network slicing with Quality of Service (QoS) to enable real-time monitoring without disrupting OT operations. Leveraging digital twin technology and Network Function Virtualization (NFV), the architecture supports predictive analytics and Industry 4.0 requirements. FL facilitates decentralized model training, preserving data privacy and scalability, though it introduces potential throughput constraints. Network slicing addresses this by creating dedicated virtualized segments optimized for performance and security. Advanced fault tolerance at the container and instance levels enhances system reliability. The proposed architecture ensures high throughput, low latency, and secure orchestration for real-time anomaly detection in OT networks. Performance evaluations validate its efficiency in throughput, deployment, and learning accuracy, providing a robust foundation for future ICS automation and data-driven decision-making.

Delgado, Brian G. Rodiles [University of Texas at ↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Towards 5G-Enabled Operational Technology for Process Monitoring and Network Slicing

Cyber-Physical Systems (CPS) are deployed to monitor physical processes in critical cyber-enabled services like power generation. However, CPS ecosystems are typically designed without robust security. While it is important to ensure optimal performance of the Operational Technology (OT) environments, security cannot be overlooked. To modernize traditional OT services, 5G technology is being integrated. 5G technology offers low latency and high availability, making it a suitable infrastructure for managing and monitoring physical processes. How-ever, integrating 5G mechanisms into large-scale OT networks introduces new implementation and performance challenges. Therefore, this paper presents a 5G-enabled CPS architecture (5G-CPS) that describes the necessary components, services, and communication protocols and conducts feasibility study to integrate 5G technology in industrial control system networks to understand the performance merits. The 5G-CPS architecture aims to minimize implementation and operational challenges associated with integrating 5G technology into constrained OT.

Aguayo, Jared M.↗

Human Mars Surface Science Operations

Human missions to the surface of Mars will have challenging science operations. This paper will explore some of those challenges, based on science operations considerations as part of more general operational concepts being developed by NASA's Human Spaceflight Architecture (HAT) Mars Destination Operations Team (DOT). The HAT Mars DOT has been developing comprehensive surface operations concepts with an initial emphasis on a multi-phased mission that includes a 500-day surface stay. This paper will address crew science activities, operational details and potential architectural and system implications in the areas of (a) traverse planning and execution, (b) sample acquisition and sample handling, (c) in-situ science analysis, and (d) planetary protection. Three cross-cutting themes will also be explored in this paper: (a) contamination control, (b) low-latency telerobotic science, and (c) crew autonomy. The present traverses under consideration are based on the report, Planning for the Scientific Exploration of Mars by Humans1, by the Mars Exploration Planning and Analysis Group (MEPAG) Human Exploration of Mars-Science Analysis Group (HEM-SAG). The traverses are ambitious and the role of science in those traverses is a key component that will be discussed in this paper. The process of obtaining, handling, and analyzing samples will be an important part of ensuring acceptable science return. Meeting planetary protection protocols will be a key challenge and this paper will explore operational strategies and system designs to meet the challenges of planetary protection, particularly with respect to the exploration of "special regions." A significant challenge for Mars surface science operations with crew is preserving science sample integrity in what will likely be an uncertain environment. Crewed mission surface assets -- such as habitats, spacesuits, and pressurized rovers -- could be a significant source of contamination due to venting, out-gassing and cleanliness levels associated with crew presence. Low-latency telerobotic science operations has the potential to address a number of contamination control and planetary protection issues and will be explored in this paper. Crew autonomy is another key cross-cutting challenge regarding Mars surface science operations, because the communications delay between earth and Mars could as high as 20 minutes one way, likely requiring the crew to perform many science tasks without direct timely intervention from ground support on earth. Striking the operational balance between crew autonomy and earth support will be a key challenge that this paper will address.

Bobskill, Marianne R.↗

New trends in photonic switching and optical networking architectures for data centers and computing systems [Invited]

The rapid increases in data traffic coupled with user preferences are driving the data center and computing system service providers to offer energy-efficient, intelligent, flexible, cost-effective, high-capacity, and low-latency data services without added complexity to the users. Disaggregated heterogeneous reconfigurable computing systems realized by photonic switching and interconnects can enhance throughput and energy efficiency for artificial intelligence/machine learning (AI/ML) workloads, especially when aided by the AI/ML-enhanced control plane. Photonic switching and new optical networking architectures are expected to solve many of these challenging problems. This paper discusses new trends in photonic switching and optical network architectures for future data centers and computing systems summarized as follows: (1) flat reconfigurable disaggregated computing enabled by high-radix photonic switching and interconnects in data centers; (2) chiplet-based computing architectures empowered by embedded photonics toward heterogeneous reconfigurable computing; (3) nanosecond-scale photonic switching in data centers and computing systems; (4) AI/ML in self-driving, application-aware, and situation-aware data centers; (5) the emergence of flexible networking for cloud computing, edge computing, and split computing, as well as flexible networking for 5G/6G RF-optical networks; and (6) the deployment of embedded co-designed silicon photonics being considered for future data centers.

Yoo, S. J. Ben (ORCID:0000000274201871)↗

Using Orbiting Carbon Observatory-2 (OCO-2) column CO2 retrievals to rapidly detect and estimate biospheric surface carbon flux anomalies

The global carbon cycle is experiencing continued perturbations via increases in atmospheric carbon concentrations, which are partly reduced by terrestrial biosphere and ocean carbon uptake. Greenhouse gas satellites have been shown to be useful in retrieving atmospheric carbon concentrations and observing surface and atmospheric CO2 seasonal-to-interannual variations. However, limited attention has been placed on using satellite column CO2 retrievals to evaluate surface CO2 fluxes from the terrestrial biosphere without advanced inversion models at low latency. Such applications could be useful to monitor, in near real time, biosphere carbon fluxes during climatic anomalies like drought, heatwaves, and floods, before more complex terrestrial biosphere model outputs and/or advanced inversion modelling estimates become available. Here, we explore the ability of Orbiting Carbon Observatory-2 (OCO-2) column-averaged dry air CO2 (XCO2) retrievals to directly detect and estimate terrestrial biosphere CO2 flux anomalies using a simple mass-balance approach. An initial global analysis of surface–atmospheric CO2 coupling and transport conditions reveals that the western US, among a handful of other regions, is a feasible candidate for using XCO2 for detecting terrestrial biosphere CO2 flux anomalies. Using the CarbonTracker model reanalysis as a test bed, we first demonstrate that a well-established mass-balance approach can estimate monthly surface CO2 flux anomalies from XCO2 enhancements in the western United States. The method is optimal when the study domain is spatially extensive enough to account for atmospheric mixing and has favorable advection conditions with contributions primarily from one background region. We find that errors in individual soundings reduce the ability of OCO-2 XCO2 to estimate more frequent, smaller surface CO2 flux anomalies. However, we find that OCO-2 XCO2 can often detect and estimate large surface flux anomalies that leave an imprint on the atmospheric CO2 concentration anomalies beyond the retrieval error/uncertainty associated with the observations. OCO-2 can thus be useful for low-latency monitoring of the monthly timing and magnitude of extreme regional terrestrial biosphere carbon anomalies.

Andrew F. Feldman↗

GAHLS: an optimized graph analytics based high level synthesis framework

The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities. With these goals in mind, we propose a novel comprehensive graph analytics based high level synthesis (GAHLS) framework that efficiently analyzes complex high level programs through a combined compiler-based approach and graph theoretic optimization and synthesizes them into message passing domain-specific accelerators. This GAHLS framework first constructs a compiler-assisted dependency graph (CaDG) from low level virtual machine (LLVM) intermediate representation (IR) of high level programs and converts it into a hardware friendly description representation. Next, the GAHLS framework performs a memory design space exploration while account for the identified computational properties from the CaDG and optimizing the system performance for higher bandwidth. The GAHLS framework also performs a robust optimization to identify the CaDG subgraphs with similar computational structures and aggregate them into intelligent processing clusters in order to optimize the usage of underlying hardware resources. Finally, the GAHLS framework synthesizes this compressed specialized CaDG into processing elements while optimizing the system performance and area metrics. Evaluations of the GAHLS framework on several real-life applications (e.g., deep learning, brain machine interfaces) demonstrate that it provides 14.27× performance improvements compared to state-of-the-art approaches such as LegUp 6.2.

97 MATHEMATICS AND COMPUTING↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗

Prioritized LT Codes

The original Luby Transform (LT) coding scheme is extended to account for data transmissions where some information symbols in a message block are more important than others. Prioritized LT codes provide unequal error protection (UEP) of data on an erasure channel by modifying the original LT encoder. The prioritized algorithm improves high-priority data protection without penalizing low-priority data recovery. Moreover, low-latency decoding is also obtained for high-priority data due to fast encoding. Prioritized LT codes only require a slight change in the original encoding algorithm, and no changes at all at the decoder. Hence, with a small complexity increase in the LT encoder, an improved UEP and low-decoding latency performance for high-priority data can be achieved. LT encoding partitions a data stream into fixed-sized message blocks each with a constant number of information symbols. To generate a code symbol from the information symbols in a message, the Robust-Soliton probability distribution is first applied in order to determine the number of information symbols to be used to compute the code symbol. Then, the specific information symbols are chosen uniform randomly from the message block. Finally, the selected information symbols are XORed to form the code symbol. The Prioritized LT code construction includes an additional restriction that code symbols formed by a relatively small number of XORed information symbols select some of these information symbols from the pool of high-priority data. Once high-priority data are fully covered, encoding continues with the conventional LT approach where code symbols are generated by selecting information symbols from the entire message block including all different priorities. Therefore, if code symbols derived from high-priority data experience an unusual high number of erasures, Prioritized LT codes can still reliably recover both high- and low-priority data. This hybrid approach decides not only "how to encode" but also "what to encode" to achieve UEP. Another advantage of the priority encoding process is that the majority of high-priority data can be decoded sooner since only a small number of code symbols are required to reconstruct high-priority data. This approach increases the likelihood that high-priority data is decoded first over low-priority data. The Prioritized LT code scheme achieves an improvement in high-priority data decoding performance as well as overall information recovery without penalizing the decoding of low-priority data, assuming high-priority data is no more than half of a message block. The cost is in the additional complexity required in the encoder. If extra computation resource is available at the transmitter, image, voice, and video transmission quality in terrestrial and space communications can benefit from accurate use of redundancy in protecting data with varying priorities.

Woo, Simon S.↗

Interferometer real time control development for SIM

This paper provides an overview of the architecture, design, integration, and test of the SIM flight interferometer real time control to meet challenging flight system requirements for the high processor throughput, low-latency interconnect, and precise synchronization to support microarcsecond-level astrometric measurements for greater than five years at 1 AU in Earth-trailing orbit.

real↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗

A Multi-Satellite Framework to Rapidly Evaluate Extreme Biosphere Cascades: The Western US 2021 Drought and Heatwave

The increasing frequency and intensity of climate extremes and complex ecosystem responses motivate the need for integrated observational studies at low-latency to determine biosphere responses and carbon-climate feedbacks. Here, we develop a satellite-based rapid attribution workflow and demonstrate its use at a 1–2-month latency to attribute drivers of the carbon cycle feedbacks during the 2020-2021 Western US drought and heatwave. In the first half of 2021, concurrent negative photosynthesis anomalies and large positive column CO 2 anomalies were detected with satellites. Using a simple atmospheric mass balance approach, we estimate a surface carbon efflux anomaly of 132 TgC in June 2021, a magnitude corroborated 28 independently with a dynamic global vegetation model. Integrated satellite observations of hydrologic processes, representing the soil-plant-atmosphere continuum (SPAC), show that these surface carbon flux anomalies are largely due to substantial reductions in photosynthesis because of a spatially widespread moisture-deficit propagation through the SPAC between 2020 and 2021. A causal model indicates deep soil moisture stores partially drove photosynthesis, maintaining its values in 2020 and driving its declines throughout 2021. The causal model also suggests legacy effects may have amplified photosynthesis deficits in 2021 beyond the direct effects of environmental forcing. The integrated, observation framework presented here provides a valuable first assessment of a biosphere extreme response and an independent testbed for improving drought propagation and mechanisms in models. The rapid identification of extreme carbon anomalies and hotspots can also aid mitigation and adaptation decisions.

Causal model↗

Air Traffic Management TestBed: Messaging Performance

The Air Traffic Management (ATM) TestBed is an air traffic management modeling and simulation platform and framework developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The communication middleware, implemented in the TestBed framework layer, is a core feature for data message exchange. The feature provides an abstraction layer called Messaging Support to allow switching one middleware to another without a need to rebuild the simulation components. Messaging performance such as latencies, run durations, and throughputs are important factors. Low latencies can produce accurate results in high-fidelity and visualization models. Short run durations are preferred because better run efficiency can be achieved. High throughputs allow more runs to be executed concurrently. This technical memorandum studies and compares the messaging performance by running a full-day, fast-time simulation using three communication middleware as well as tweaking the default communication middleware settings used by the TestBed. Results indicate that the messaging performance could be improved by disabling either compression or persistence settings, while the run duration and throughput could be further improved by disabling both settings with a tradeoff of the message latencies increased by a factor of ten.

Chok Fung Lai↗

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING↗

Short-Block Protograph-Based LDPC Codes

Short-block low-density parity-check (LDPC) codes of a special type are intended to be especially well suited for potential applications that include transmission of command and control data, cellular telephony, data communications in wireless local area networks, and satellite data communications. [In general, LDPC codes belong to a class of error-correcting codes suitable for use in a variety of wireless data-communication systems that include noisy channels.] The codes of the present special type exhibit low error floors, low bit and frame error rates, and low latency (in comparison with related prior codes). These codes also achieve low maximum rate of undetected errors over all signal-to-noise ratios, without requiring the use of cyclic redundancy checks, which would significantly increase the overhead for short blocks. These codes have protograph representations; this is advantageous in that, for reasons that exceed the scope of this article, the applicability of protograph representations makes it possible to design highspeed iterative decoders that utilize belief- propagation algorithms.

Divsalar, Dariush↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

PERSIANN-Unet: A Global Deep Learning Framework for Near-Real-Time Precipitation Estimation Using Infrared Data

Access to high-quality, high-resolution, near-real-time precipitation data is essential for hydrological and meteorological research and disaster mitigation. Traditional tools such as rain gauges and radar networks, though effective, have limitations, including sparse coverage in remote areas and high operational costs. Satellite data, with its global coverage and high spatial and temporal resolutions, mitigates limitations in coverage. Satellite precipitation products like Hydro Estimator (HE), Integrated Multi-satellitE Retrievals for Global Precipitation Measurement (IMERG), and Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks (PERSIANN) utilize both geosynchronous thermal infrared (IR) and passive microwave (PMW) data in their operation. PMW sensors offer detailed atmospheric profiles but suffer from higher latency, whereas IR sensors provide lower latency but only capture cloud-top information. Despite this constraint, IR data remains attractive for low-latency precipitation estimation. Recent advances in deep learning, particularly convolutional neural networks (CNNs), have further improved satellite precipitation retrievals. This study introduces PERSIANN-Unet (PUnet or PERSIANN V3), a quasi-global algorithm covering 60°N–60°S that combines IR data, monthly climatology, and the UNet architecture to produce half-hourly precipitation estimates at 0.04° resolution. The product is evaluated against HE, IMERG, and PDIR-Now for 2022–2023. Results show that PUnet closely matches its training target, IMERG V07 Final, at the global scale, and performance is further evaluated against Stage IV as a reference over CONUS. Training PUnet on IMERG (2016–2021) leverages a high-quality, integrated PMW IR-gauge precipitation product while developing an IR-based framework not reliant on PMW availability. By operating on a single global image, PUnet avoids tile partitioning and blending steps, reducing edge discontinuities, and produces more spatially consistent precipitation fields across hemispheres.

Phu Nguyen↗

Challenges Using the Linux Network Stack for Real-Time Communication

Starting in the early 2000s, human-in-the-loop (HITL) simulation groups at NASA and the Air Force Research Lab began using the Linux network stack for some real-time communication. More recently, SpaceX has adopted Ethernet as the primary bus technology for its Falcon launch vehicles and Dragon capsules. As the Linux network stack makes its way from ground facilities to flight critical systems, it is necessary to recognize that the network stack is optimized for communication over the open Internet, which cannot provide latency guarantees. The Internet protocols and their implementation in the Linux network stack contain numerous design decisions that favor throughput over determinism and latency. These decisions often require workarounds in the application or customization of the stack to maintain a high probability of low latency on closed networks, especially if the network must be fault tolerant to single event upsets.

Madden, Michael M.↗