Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Towards Fully Secure 5G Ultra-Low Latency Communications: A Cost-Security Functions Analysis

Future components to enhance the basic, native security of 5G networks are either complex mechanisms whose impact in the requiring 5G communications are not considered, or lightweight solutions adapted to ultra-reliable low-latency communications (URLLC) but whose security properties remain under discussion. Although different 5G network slices may have different requirements, in general, both visions seem to fall short at provisioning secure URLLC in the future. In this work we address this challenge, by introducing cost-security functions as a method to evaluate the performance and adequacy of most developed and employed non-native enhanced security mechanisms in 5G networks. We categorize those new security components into different groups according to their purpose and deployment scope. We propose to analyze them in the context of existing 5G architectures using two different approaches. First, using model checking techniques, we will evaluate the probability of an attacker to be successful against each security solution. Second, using analytical models, we will analyze the impact of these security mechanisms in terms of delay, throughput consumption, and reliability. Finally, we will combine both approaches using stochastic cost-security functions and the PRISM model checker to create a global picture. Our results are first evidence of how a 5G network that covers and strengthened all security areas through enhanced, dedicated non-native mechanisms could only guarantee secure URLLC with a probability of ~55%.

5G networks↗

A low-latency graph computer to identify metastable particles at the Large Hadron Collider for real-time analysis of potential dark matter signatures

Abstract Image recognition is a pervasive task in many information-processing environments. We present a solution to a difficult pattern recognition problem that lies at the heart of experimental particle physics. Future experiments with very high-intensity beams will produce a spray of thousands of particles in each beam-target or beam-beam collision. Recognizing the trajectories of these particles as they traverse layers of electronic sensors is a massive image recognition task that has never been accomplished in real time. We present a real-time processing solution that is implemented in a commercial field-programmable gate array using high-level synthesis. It is an unsupervised learning algorithm that uses techniques of graph computing. A prime application is the low-latency analysis of dark-matter signatures involving metastable charged particles that manifest as disappearing tracks.

47 OTHER INSTRUMENTATION↗

Ultra-low latency recurrent neural network inference on FPGAs for physics applications with hls4ml

Abstract Recurrent neural networks have been shown to be effective architectures for many tasks in high energy physics, and thus have been widely adopted. Their use in low-latency environments has, however, been limited as a result of the difficulties of implementing recurrent architectures on field-programmable gate arrays (FPGAs). In this paper we present an implementation of two types of recurrent neural network layers—long short-term memory and gated recurrent unit—within the hls4ml framework. We demonstrate that our implementation is capable of producing effective designs for both small and large models, and can be customized to meet specific design requirements for inference latencies and FPGA resources. We show the performance and synthesized designs for multiple neural networks, many of which are trained specifically for jet identification tasks at the CERN Large Hadron Collider.

97 MATHEMATICS AND COMPUTING↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

Low Latency and High Data Rate (LLHD) Scheduler: A Multipath TCP Scheduler for Dynamic and Heterogeneous Networks

The scheduler is a crucial component of the multipath transmission control protocol (MPTCP) that dictates the path that a data packet takes. Schedulers are in charge of delivering data packets in the right order to prevent delays caused by head-of-line blocking. The modern Internet is a complicated network whose characteristics change in real-time. MPTCP schedulers are supposed to understand the real-time properties of the underlying network, such as latency, path loss, and capacity, in order to make appropriate scheduling decisions. However, the present scheduler does not take into account all of these characteristics together, resulting in lower performance. We present the low latency and high data rate (LLHD) scheduler, which successfully makes scheduling decisions based on real-time information on latency, path loss, and capacity, and achieves around 25% higher throughput and 45% lower data transmission delay than Linux’s default MPTCP scheduler.

97 MATHEMATICS AND COMPUTING↗

Low-latency NuMI Trigger for the CHIPS-5 Neutrino Detector

The CHIPS R&D project aims to develop affordable water Cherenkov detectors for large-scale underwater installations. In 2019, a 5kt prototype detector CHIPS-5 was deployed in northern Minnesota to study neutrinos generated by the nearby NuMI beam. This contribution presents a dedicated low-latency time distribution system for CHIPS-5 that delivers timing signals from the Fermilab accelerator to the detector with sub-nanosecond precision. Exploiting existing NOvA infrastructure, the time distribution system achieves this only with open-source software and conventional network elements. In a time-of-flight study, the presented system has reliably offered a time budget of $610 \pm 330\text{ ms}$ for on-site triggering. This permits advanced analysis in real-time as well as a novel hardware-assisted active triggering mode, which reduces DAQ computing load and network bandwidth outside triggered time windows.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Ps and Qs: Quantization-Aware Pruning for Efficient Low Latency Neural Network Inference

Efficient machine learning implementations optimized for inference in hardware have wide-ranging benefits, depending on the application, from lower inference latency to higher data throughput and reduced energy consumption. Two popular techniques for reducing computation in neural networks are pruning, removing insignificant synapses, and quantization, reducing the precision of the calculations. In this work, we explore the interplay between pruning and quantization during the training of neural networks for ultra low latency applications targeting high energy physics use cases. Techniques developed for this study have potential applications across many other domains. We study various configurations of pruning during quantization-aware training, which we term quantization-aware pruning, and the effect of techniques like regularization, batch normalization, and different pruning schemes on performance, computational complexity, and information content metrics. We find that quantization-aware pruning yields more computationally efficient models than either pruning or quantization alone for our task. Further, quantization-aware pruning typically performs similar to or better in terms of computational efficiency compared to other neural architecture search techniques like Bayesian optimization. Surprisingly, while networks with different training configurations can have similar performance for the benchmark application, the information content in the network can vary significantly, affecting its generalizability.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

App2Net: Moving App Functions to Network & a Case Study on Low-latency Feedback

Recent advances in programmable networks enable custom processing of data at hundreds of gigabits per second. These advances can boost the performance of many distributed applications. Yet the high-level languages used by application developers are different from the data plane programming languages (such as P4 and NPL) used by network equipment. This language barrier slows innovation. Our hourglass-shaped architectural solution aims to lower this language barrier. This enables the application developer community to leverage programmable networks for achieving better performance. In this paper we propose a JSON-based intermediate representation to bridge the gap between applications and in-network computing. We demonstrate an instance of the solution in the context of a low-latency feedback application that enables SQL-based data filtering in a P4-based programmable environment. We also present a prototype compiler to convert an intermediate representation in JSON to P4 source.

Sankaran, Ganesh↗

Towards On-Chip Learning for Low Latency Reasoning with End-to-End Synthesis

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application-Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frame-works), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This paper provides an overview of the flow in the context of the generation of accelerators for edge processing to be integrated in transmission electron microscopy (TEM) devices, focusing on use cases from precision material synthesis. We show the tool in action with an example of design space exploration for inference on reconfigurable devices with a conventional deep neural network model (LeNet). Finally, we discuss the research directions and opportunities enabled by SODA in the area of autonomous control for scientific experimental workflows.

Castellana, Vito G.↗

Low latency optical-based mode tracking with machine learning deployed on FPGAs on a tokamak

Active feedback control in magnetic confinement fusion devices is desirable to mitigate plasma instabilities and enable robust operation. Optical high-speed cameras provide a powerful, non-invasive diagnostic and can be suitable for these applications. Here, in this study, we process high-speed camera data, at rates exceeding 100 kfps, on in situ field-programmable gate array (FPGA) hardware to track magnetohydrodynamic (MHD) mode evolution and generate control signals in real time. Our system utilizes a convolutional neural network (CNN) model, which predicts the n = 1 MHD mode amplitude and phase using camera images with better accuracy than other tested non-deep-learning-based methods. By implementing this model directly within the standard FPGA readout hardware of the high-speed camera diagnostic, our mode tracking system achieves a total trigger-to-output latency of 17.6 μs and a throughput of up to 120 kfps. This study at the High Beta Tokamak-Extended Pulse (HBT-EP) experiment demonstrates an FPGA-based high-speed camera data acquisition and processing system, enabling application in real-time machine-learning-based tokamak diagnostic and control as well as potential applications in other scientific domains.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Low-latency quantum control using AI algorithms on cryogenic microelectronics platforms

Superconducting quantum computers are a particular technology that involve the realization of qubits by means of LC circuits with a Josephson junction. As shown in Fig. 1, this latter addition has the effect of introducing a nonlinearity (anharmonicity) that makes the frequency spacing between each energy level non-uniform. This is a desired effect for separating the two lowest levels, used for computation, from the higher ones, that are theoretically infinite and which are not desired to be used.

97 MATHEMATICS AND COMPUTING↗

Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications

With more scientific fields relying on neural networks (NNs) to process data incoming at extreme throughputs and latencies, it is crucial to develop NNs with all their parameters stored on-chip. In many of these applications, there is not enough time to go off-chip and retrieve weights. Even more so, off-chip memory such as DRAM does not have the bandwidth required to process these NNs as fast as the data is being produced (e.g., every 25 ns). As such, these extreme latency and bandwidth requirements have architectural implications for the hardware intended to run these NNs: 1) all NN parameters must fit on-chip, and 2) codesigning custom/reconfigurable logic is often required to meet these latency and bandwidth constraints. In our work, we show that many scientific NN applications must run fully on chip, in the extreme case requiring a custom chip to meet such stringent constraints.

Weng, Olivia↗