Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The Impact of Traffic Prioritization on Deep Space Network Mission Traffic

A select number of missions supported by NASA's Deep Space Network (DSN) are demanding very high data rates. For example, the Kepler Mission was launched March 7, 2009 and at that time required the highest data rate of any NASA mission, with maximum rates of 4.33 Mb/s being provided via Ka band downlinks. The James Webb Space Telescope will require a maximum 28 Mb/s science downlink data rate also using Ka band links; as of this writing the launch is scheduled for a June 2014 launch. The Lunar Reconnaissance Orbiter, launched June 18, 2009, has demonstrated data rates at 100 Mb/s at lunar-Earth distances using NASA's Near Earth Network (NEN) and K-band. As further advances are made in high data rate space telecommunications, particularly with emerging optical systems, it is expected that large surges in demand on the supporting ground systems will ensue. A performance analysis of the impact of high variance in demand has been conducted using our Multi-mission Advanced Communications Hybrid Environment for Test and Evaluation (MACHETE) simulation tool. A comparison is made regarding the incorporation of Quality of Service (QoS) mechanisms and the resulting ground-to-ground Wide Area Network (WAN) bandwidth necessary to meet latency requirements across different user missions. It is shown that substantial reduction in WAN bandwidth may be realized through QoS techniques when low data rate users with low-latency needs are mixed with high data rate users having delay-tolerant traffic.

prioritization↗

The Impact of Traffic Prioritization on Deep Space Network Mission Traffic

A select number of missions supported by NASA's Deep Space Network (DSN) are demanding very high data rates. For example, the Kepler Mission was launched March 7, 2009 and at that time required the highest data rate of any NASA mission, with maximum rates of 4.33 Mb/s being provided via Ka band downlinks. The James Webb Space Telescope will require a maximum 28 Mb/s science downlink data rate also using Ka band links; as of this writing the launch is scheduled for a June 2014 launch. The Lunar Reconnaissance Orbiter, launched June 18, 2009, has demonstrated data rates at 100 Mb/s at lunar-Earth distances using NASA's Near Earth Network (NEN) and K-band. As further advances are made in high data rate space telecommunications, particularly with emerging optical systems, it is expected that large surges in demand on the supporting ground systems will ensue. A performance analysis of the impact of high variance in demand has been conducted using our Multi-mission Advanced Communications Hybrid Environment for Test and Evaluation (MACHETE) simulation tool. A comparison is made regarding the incorporation of Quality of Service (QoS) mechanisms and the resulting ground-to-ground Wide Area Network (WAN) bandwidth necessary to meet latency requirements across different user missions. It is shown that substantial reduction in WAN bandwidth may be realized through QoS techniques when low data rate users with low-latency needs are mixed with high data rate users having delay-tolerant traffic.

quality of service↗

Nanosecond anomaly detection with decision trees and real-time application to exotic Higgs decays

Abstract We present an interpretable implementation of the autoencoding algorithm, used as an anomaly detector, built with a forest of deep decision trees on FPGA, field programmable gate arrays. Scenarios at the Large Hadron Collider at CERN are considered, for which the autoencoder is trained using known physical processes of the Standard Model. The design is then deployed in real-time trigger systems for anomaly detection of unknown physical processes, such as the detection of rare exotic decays of the Higgs boson. The inference is made with a latency value of 30 ns at percent-level resource usage using the Xilinx Virtex UltraScale+ VU9P FPGA. Our method offers anomaly detection at low latency values for edge AI users with resource constraints.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dual Purpose Simulation: New Data Link Test and Comparison With VDL-2

While the results of this paper are similar to those of previous research, in this paper technical difficulties present there are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL- 2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PCSMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA outperforms traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. We are testing a new and better data link that could replace CSMA with relative ease. Work is underway to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

Dual Purpose Simulation: New Data Link Test and Performance Limit Testing of Currently Deployed Data Link

While the results of this paper are similar to those of [I], in this paper technical difficulties present in [I] are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air-traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft-taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PCSMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA outperforms traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. So we have the tools to determine the traffic-loading conditions where traditional CSMA will fail, and we are testing a new and better data link that could replace it with relative ease. Work is currently being done to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

Dual Purpose Simulation: New Data Link Test and Comparison with VDL-2

While the results of this paper are similar to those of previous research, in this paper technical difficulties present there are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (A TN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft- taking off, flying realistic free- flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PC SMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA out performs traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. We are testing a new and better data link that could replace CSMA with relative ease. Work is underway to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

CSMA Versus Prioritized CSMA for Air-Traffic-Control Improvement

OPNET version 7.0 simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link, Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air-traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. There are 32 airports in the simulation, 29 of which are either sources or destinations for the air-traffic of the aforementioned three airports. The simulation involves 111 Air Traffic Control (ATC) ground stations, and 1,235 equally equipped aircraft-taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collisionless, Prioritized Carrier Sense Multiple Access (CSMA) is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, Prioritized CSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of Prioritized CSMA for implementing low latency, high throughput, and efficient connectivity.

Robinson, Daryl C.↗

Efficient Routing of Quantum LDPC Codes on Programmable 2D Toric Architectures

Quantum low-density parity-check codes are promising candidates towards scalable fault-tolerant quantum computation. Among these, bivariate bicycle (BB) codes offer superior encoding rates and large code distance compared to surface codes. However, their requirement on long-range stabilizer measurements poses significant challenges for implementation on realistic hardware with limited connectivity, such as superconducting circuit platforms. In this work, we introduce a novel hardware-software co-design that leverages a programmable communication network architecture to address these limitations. Our approach utilizes a 2D toric network of oscillators as a flexible communication fabric linking qubits at each site. Such architecture significantly reduces the number of long-range couplers required from O ( n ) to O (√ n ). Dual-rail qubits, along with native gates including Swap-Wait-Swap gates and beamsplitter SWAPs, ensure that long-range two-qubit gates can be executed with high fidelity and low latency. To further enhance performance, our qubit layout and routing algorithm utilize symmetries of the codes and enable maximum parallelism for long-range two-qubit gates, maintaining a low syndrome extraction cycle duration and scalability over the code length. We perform circuit-level simulation with realistic noise modeling based on experimental hardware parameters, observing an logical error rate per logical qubit per cycle of 3.06% for [[18,4,4]] BB code, 2.6× less than the existing experimental result. These findings provide a practical roadmap and identify key technological advancements needed to achieve low-overhead fault-tolerant quantum computing at scale.

Liu, Kun [Yale Univ., New Haven, CT (United States↗

HTMT-class Latency Tolerant Parallel Architecture for Petaflops Scale Computation

Computational Aero Sciences and other numeric intensive computation disciplines demand computing throughputs substantially greater than the Teraflops scale systems only now becoming available. The related fields of fluids, structures, thermal, combustion, and dynamic controls are among the interdisciplinary areas that in combination with sufficient resolution and advanced adaptive techniques may force performance requirements towards Petaflops. This will be especially true for compute intensive models such as Navier-Stokes are or when such system models are only part of a larger design optimization computation involving many design points. Yet recent experience with conventional MPP configurations comprising commodity processing and memory components has shown that larger scale frequently results in higher programming difficulty and lower system efficiency. While important advances in system software and algorithms techniques have had some impact on efficiency and programmability for certain classes of problems, in general it is unlikely that software alone will resolve the challenges to higher scalability. As in the past, future generations of high-end computers may require a combination of hardware architecture and system software advances to enable efficient operation at a Petaflops level. The NASA led HTMT project has engaged the talents of a broad interdisciplinary team to develop a new strategy in high-end system architecture to deliver petaflops scale computing in the 2004/5 timeframe. The Hybrid-Technology, MultiThreaded parallel computer architecture incorporates several advanced technologies in combination with an innovative dynamic adaptive scheduling mechanism to provide unprecedented performance and efficiency within practical constraints of cost, complexity, and power consumption. The emerging superconductor Rapid Single Flux Quantum electronics can operate at 100 GHz (the record is 770 GHz) and one percent of the power required by convention semiconductor logic. Wave Division Multiplexing optical communications can approach a peak per fiber bandwidth of 1 Tbps and the new Data Vortex network topology employing this technology can connect tens of thousands of ports providing a bi-section bandwidth on the order of a Petabyte per second with latencies well below 100 nanoseconds, even under heavy loads. Processor-in-Memory (PIM) technology combines logic and memory on the same chip exposing the internal bandwidth of the memory row buffers at low latency. And holographic storage photorefractive storage technologies provide high-density memory with access a thousand times faster than conventional disk technologies. Together these technologies enable a new class of shared memory system architecture with a peak performance in the range of a Petaflops but size and power requirements comparable to today's largest Teraflops scale systems. To achieve high-sustained performance, HTMT combines an advanced multithreading processor architecture with a memory-driven coarse-grained latency management strategy called "percolation", yielding high efficiency while reducing the much of the parallel programming burden. This paper will present the basic system architecture characteristics made possible through this series of advanced technologies and then give a detailed description of the new percolation approach to runtime latency management.

Sterling, Thomas↗

Protograph-Based Raptor-Like Codes

Theoretical analysis has long indicated that feedback improves the error exponent but not the capacity of pointto- point memoryless channels. The analytic and empirical results indicate that at short blocklength regime, practical rate-compatible punctured convolutional (RCPC) codes achieve low latency with the use of noiseless feedback. In 3GPP, standard rate-compatible turbo codes (RCPT) did not outperform the convolutional codes in the short blocklength regime. The reason is the convolutional codes for low number of states can be decoded optimally using Viterbi decoder. Despite excellent performance of convolutional codes at very short blocklengths, the strength of convolutional codes does not scale with the blocklength for a fixed number of states in its trellis.

Divsalar, Dariush↗

Accelerating data acquisition with FPGA-based edge machine learning: a case study with LCLS-II

New scientific experiments and instruments generate vast amounts of data that need to be transferred for storage or further processing, often overwhelming traditional systems. Edge machine learning (EdgeML) addresses this challenge by integrating machine learning (ML) algorithms with edge computing, enabling real-time data processing directly at the point of data generation. EdgeML is particularly beneficial for environments where immediate decisions are required, or where bandwidth and storage are limited. In this paper, we demonstrate a high-speed configurable ML model in a fully customizable EdgeML system using a field programmable gate array (FPGA). Our demonstration focuses on an angular array of electron spectrometers, referred to as the ‘CookieBox,’ developed for the Linac Coherent Light Source II project. The EdgeML system captures 51.2 Gbps from a 6.4 GS s −1 analog to digital converter and is designed to integrate data pre-processing and ML inside an FPGA. Our implementation achieves an inference latency of 0.2 µs for the ML model, and a total latency of 0.4 µs for the complete EdgeML system, which includes pre-processing, data transmission, digitization, and ML inference. The modular design of the system allows it to be adapted for other instrumentation applications requiring low-latency data processing.

97 MATHEMATICS AND COMPUTING↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Near Optimality of Matched Filter Detection for Cyclic Prefix Direct Sequence Spread Spectrum

Abstract—Cyclic Prefix Direct Sequence Spread Spectrum -DSSS) has been presented as a potential solution for ultrareliable low latency communications (URLLC) and massive machine type communication (mMTC), where the CP-DSSS waveform would operate as a secondary network at the same frequencies as the primary network but at much lower SNR. In this paper, we show that when operating in the low SNR regime, CP-DSSS achieves near optimum performance when using a matched filter (MF) detector at the receiver. Time reversal (TR) precoding at the transmitter is also analyzed. These results also carry forward into multi-antenna scenarios where array gain is preserved. With nearly optimal performance of MF detection, CP-DSSS can be implemented with simple device transceiver structures, reducing per-unit cost for massively deployed networks.

5G and Beyond Communications↗

NASA’s Land, Atmosphere near Real-Time Capability for EOS ( LANCE) @10 Years: A Look Back at Its Origins in MODIS Terra

This poster looks back on how the first near real-time (NRT) images from MODIS Terra provided the impetus for the creation of the Land, Atmosphere Near Real-Time Capability for EOS (LANCE) – a near real-time (NRT) capability that currently serves low latency products for monitoring air quality, floods, duststorms, snow cover and agriculture, as well as for public education and outreach to users in over 160 countries.

Davies, Diane↗

Programmable low-coherence wavefronts for enhanced localization

Engineering the properties of electromagnetic wavefronts has become essential to imaging, wireless security, sensing, and wireless communication. In particular, wavefronts that exhibit low spatial coherence can enable sensing functionalities with high accuracy and low latency. The typical use of such wavefronts cannot take advantage of these possibilities, as they require the ability to dynamically reconfigure the wavefront in a controllable and repeatable fashion, over a broad spectral bandwidth. Here, we propose a new approach for generating broadband reconfigurable wavefronts which not only exhibit low spatial coherence at a particular frequency, but are also decorrelated with the wavefronts simultaneously generated at other frequencies. We demonstrate that this frequency-domain decorrelation is a key feature that, in combination with dynamic reconfigurability, enables localization measurements with an order-of-magnitude improvement in accuracy compared to the state of the art.

36 MATERIALS SCIENCE↗

Using Transparent Informed Prefetching (TIP) to reduce file read latency

As processor performance gains continue to outstrip Input/Output gains, I/O performance is becoming critical to overall system performance. File read latency is the most significant bottleneck for high performance I/O. Other aspects of I/O performance benefit from recent advances in disk bandwidth and throughput resulting from disk arrays, and in write performance derived from buffered write behind and the Log-structured File System. The access gap problem limiting improvements in read latency is exacerbated by distributed file systems operating over networks with diverse bandwidth. Focus is on extending the power of caching and prefetching to reduce file read latencies by exploiting hints from high-levels of a system. Such Transparent Informed Prefetching, TIP, and its benefits are described. It is argued that hints that disclose high level knowledge are a means for transferring optimization information across, without violating, module boundaries. How TIP can be used to convert the high throughput of new technologies such as disk arrays and log-structured file systems into low latency for applications is discussed. Our preliminary experiments show reductions in wall - clock execution time of 13 percent and 20 percent for a multiple module compilation tool (make) accessing data on a local disk and remote Coda file server, respectively, and a reduction of 30 percent for a text search (grep) remotely accessing many small files.

Patterson, R. H.↗

Variational autoencoders for at-source data reduction and anomaly detection in high energy particle detectors

Detectors in next-generation high-energy physics experiments face several daunting requirements, such as high data rates, damaging radiation exposure, and stringent constraints on power, space, and latency. To address these challenges, machine learning in readout electronics can be leveraged for smart detector designs, enabling intelligent inference and data reduction at-source. Variational autoencoders (VAEs) offer a variety of benefits for front-end readout; an on-sensor encoder can perform efficient lossy data compression while simultaneously providing a latent space representation that can be used for anomaly detection. Results are presented from low-latency and resource-efficient VAEs for front-end data processing in a futuristic silicon pixel detector. Encoder-based data compression is found to preserve good performance of off-detector analysis while significantly reducing the off-detector data rate as compared to a similarly sized data filtering approach. Furthermore, the latent space information is found to be a useful discriminator in the context of real-time sensor defect monitoring. Together, these results highlight the multifaceted utility of autoencoder-based front-end readout schemes and motivate their consideration in future detector designs.

47 OTHER INSTRUMENTATION↗

Telerobotics Workstation (TRWS) for Deep Space Habitats

On medium- to long-duration human spaceflight missions, latency in communications from Earth could reduce efficiency or hinder local operations, control, and monitoring of the various mission vehicles and other elements. Regardless of the degree of autonomy of any one particular element, a means of monitoring and controlling the elements in real time based on mission needs would increase efficiency and response times for their operation. Since human crews would be present locally, a local means for monitoring and controlling all the various mission elements is needed, particularly for robotic elements where response to interesting scientific features in the environment might need near- instantaneous manipulation and control. One of the elements proposed for medium- and long-duration human spaceflight missions, the Deep Space Habitat (DSH), is intended to be used as a remote residence and working volume for human crews. The proposed solution for local monitoring and control would be to provide a workstation within the DSH where local crews can operate local vehicles and robotic elements with little to no latency. The Telerobotics Workstation (TRWS) is a multi-display computer workstation mounted in a dedicated location within the DSH that can be adjusted for a variety of configurations as required. From an Intra-Vehicular Activity (IVA) location, the TRWS uses the Robot Application Programming Interface Delegate (RAPID) control environment through the local network to remotely monitor and control vehicles and robotic assets located outside the pressurized volume in the immediate vicinity or at low-latency distances from the habitat. The multiple display area of the TRWS allows the crew to have numerous windows open with live video feeds, control windows, and data browsers, as well as local monitoring and control of the DSH and associated systems.

Mittman, David S.↗