Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Low latency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SatCORPS Global Cloud Composite (GCC): the Design and Delivery of A High Quality, High Resolution, Global Cloud Product Available in Near-Real Time

The NASA Satellite ClOud and Radiation Property retrieval System (SatCORPS) supports the development of an analysis ready and cloud-optimized data transformation pipeline and geospatial service enablement of a global cloud composite (GCC) product derived from global geostationary satellite imagery. This geospatial service will be available at high temporal and spatial resolution via the SatCORPS web mapping application for visualization and analysis as well as direct ingestion to common geospatial software and custom programming. The resulting global cloud composite products from the processing pipeline can then be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. Near real time global observations are created through the composition of five geostationary satellites that provides modelling and forecasting communities with the capability to provide high quality and timely information to start the projection process. The Global Cloud Composite product combines information from geostationary satellites, GOES-16, GOES-17, Himawari-8, Meteosat-11 and Meteosat-9 to create a single global composite netcdf file and images using the different products within the netcdf file. The SatCORPS team, though our Global Cloud Composite (GCC) product and web-based visualization tools including Geographical Information System (GIS) services provide near real time global cloud product information to both automated processes and traditional web users that is timely and high quality derived from geostationary satellites. The Global Cloud Composite product takes advantage of the scalable processing resources provided by the AWS batch service to provide new composites every thirty minutes. Because information from each of the low earth orbiting satellites is available on schedules tuned to the specific satellite, the processing algorithm temporally composites the final dataset as each satellite’s information becomes available. The SatCORPS team has leveraged our experience using Amazon Web Services (AWS) to build a low latency high availability tool that allows end users both human and automated to acquire high quality and high-resolution Geostationary Earth Orbiting (GEO) information at zero cost to the end user. This presentation will describe how we architected and implemented the service as well as lessons learned based on our experiences both developing and operating the system. The lessons learned include how we integrated multiple services including Amazon Batch, Amazon S3 and Amazon Lambda service to create a low cost but high-performance processing system that is capable of identifying and processing the most appropriate satellite overpass information into global cloud composites. We will also describe our web-based tools including our Geographic Information System that can be used for visualization and analysis. The products from the processing can be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. The SatCORPS Global Composite Cloud product provides sophisticated global composited cloud research products with very low latency that we see that as filling a rapidly growing need in the research and modelling community with no up-front nor ongoing costs associated with downloading or using the information.

AWS AMCE SMCE GCC SATCORPS GLOBAL CLOUD COMPOSITE ↗

An open, parallel I/O computer as the platform for high-performance, high-capacity mass storage systems

APTEC Computer Systems is a Portland, Oregon based manufacturer of I/O computers. APTEC's work in the context of high density storage media is on programs requiring real-time data capture with low latency processing and storage requirements. An example of APTEC's work in this area is the Loral/Space Telescope-Data Archival and Distribution System. This is an existing Loral AeroSys designed system, which utilizes an APTEC I/O computer. The key attributes of a system architecture that is suitable for this environment are as follows: (1) data acquisition alternatives; (2) a wide range of supported mass storage devices; (3) data processing options; (4) data availability through standard network connections; and (5) an overall system architecture (hardware and software designed for high bandwidth and low latency). APTEC's approach is outlined in this document.

Abineri, Adrian↗

Human Mars Surface Science Operations

Human missions to the surface of Mars will have challenging science operations. This paper will explore some of those challenges, based on science operations considerations as part of more general operational concepts being developed by NASA's Human Spaceflight Architecture (HAT) Mars Destination Operations Team (DOT). The HAT Mars DOT has been developing comprehensive surface operations concepts with an initial emphasis on a multi-phased mission that includes a 500-day surface stay. This paper will address crew science activities, operational details and potential architectural and system implications in the areas of (a) traverse planning and execution, (b) sample acquisition and sample handling, (c) in-situ science analysis, and (d) planetary protection. Three cross-cutting themes will also be explored in this paper: (a) contamination control, (b) low-latency telerobotic science, and (c) crew autonomy. The present traverses under consideration are based on the report, Planning for the Scientific Exploration of Mars by Humans1, by the Mars Exploration Planning and Analysis Group (MEPAG) Human Exploration of Mars-Science Analysis Group (HEM-SAG). The traverses are ambitious and the role of science in those traverses is a key component that will be discussed in this paper. The process of obtaining, handling, and analyzing samples will be an important part of ensuring acceptable science return. Meeting planetary protection protocols will be a key challenge and this paper will explore operational strategies and system designs to meet the challenges of planetary protection, particularly with respect to the exploration of "special regions." A significant challenge for Mars surface science operations with crew is preserving science sample integrity in what will likely be an uncertain environment. Crewed mission surface assets -- such as habitats, spacesuits, and pressurized rovers -- could be a significant source of contamination due to venting, out-gassing and cleanliness levels associated with crew presence. Low-latency telerobotic science operations has the potential to address a number of contamination control and planetary protection issues and will be explored in this paper. Crew autonomy is another key cross-cutting challenge regarding Mars surface science operations, because the communications delay between earth and Mars could as high as 20 minutes one way, likely requiring the crew to perform many science tasks without direct timely intervention from ground support on earth. Striking the operational balance between crew autonomy and earth support will be a key challenge that this paper will address.

Bobskill, Marianne R.↗

Using Orbiting Carbon Observatory-2 (OCO-2) column CO2 retrievals to rapidly detect and estimate biospheric surface carbon flux anomalies

The global carbon cycle is experiencing continued perturbations via increases in atmospheric carbon concentrations, which are partly reduced by terrestrial biosphere and ocean carbon uptake. Greenhouse gas satellites have been shown to be useful in retrieving atmospheric carbon concentrations and observing surface and atmospheric CO2 seasonal-to-interannual variations. However, limited attention has been placed on using satellite column CO2 retrievals to evaluate surface CO2 fluxes from the terrestrial biosphere without advanced inversion models at low latency. Such applications could be useful to monitor, in near real time, biosphere carbon fluxes during climatic anomalies like drought, heatwaves, and floods, before more complex terrestrial biosphere model outputs and/or advanced inversion modelling estimates become available. Here, we explore the ability of Orbiting Carbon Observatory-2 (OCO-2) column-averaged dry air CO2 (XCO2) retrievals to directly detect and estimate terrestrial biosphere CO2 flux anomalies using a simple mass-balance approach. An initial global analysis of surface–atmospheric CO2 coupling and transport conditions reveals that the western US, among a handful of other regions, is a feasible candidate for using XCO2 for detecting terrestrial biosphere CO2 flux anomalies. Using the CarbonTracker model reanalysis as a test bed, we first demonstrate that a well-established mass-balance approach can estimate monthly surface CO2 flux anomalies from XCO2 enhancements in the western United States. The method is optimal when the study domain is spatially extensive enough to account for atmospheric mixing and has favorable advection conditions with contributions primarily from one background region. We find that errors in individual soundings reduce the ability of OCO-2 XCO2 to estimate more frequent, smaller surface CO2 flux anomalies. However, we find that OCO-2 XCO2 can often detect and estimate large surface flux anomalies that leave an imprint on the atmospheric CO2 concentration anomalies beyond the retrieval error/uncertainty associated with the observations. OCO-2 can thus be useful for low-latency monitoring of the monthly timing and magnitude of extreme regional terrestrial biosphere carbon anomalies.

Andrew F. Feldman↗

Prioritized LT Codes

The original Luby Transform (LT) coding scheme is extended to account for data transmissions where some information symbols in a message block are more important than others. Prioritized LT codes provide unequal error protection (UEP) of data on an erasure channel by modifying the original LT encoder. The prioritized algorithm improves high-priority data protection without penalizing low-priority data recovery. Moreover, low-latency decoding is also obtained for high-priority data due to fast encoding. Prioritized LT codes only require a slight change in the original encoding algorithm, and no changes at all at the decoder. Hence, with a small complexity increase in the LT encoder, an improved UEP and low-decoding latency performance for high-priority data can be achieved. LT encoding partitions a data stream into fixed-sized message blocks each with a constant number of information symbols. To generate a code symbol from the information symbols in a message, the Robust-Soliton probability distribution is first applied in order to determine the number of information symbols to be used to compute the code symbol. Then, the specific information symbols are chosen uniform randomly from the message block. Finally, the selected information symbols are XORed to form the code symbol. The Prioritized LT code construction includes an additional restriction that code symbols formed by a relatively small number of XORed information symbols select some of these information symbols from the pool of high-priority data. Once high-priority data are fully covered, encoding continues with the conventional LT approach where code symbols are generated by selecting information symbols from the entire message block including all different priorities. Therefore, if code symbols derived from high-priority data experience an unusual high number of erasures, Prioritized LT codes can still reliably recover both high- and low-priority data. This hybrid approach decides not only "how to encode" but also "what to encode" to achieve UEP. Another advantage of the priority encoding process is that the majority of high-priority data can be decoded sooner since only a small number of code symbols are required to reconstruct high-priority data. This approach increases the likelihood that high-priority data is decoded first over low-priority data. The Prioritized LT code scheme achieves an improvement in high-priority data decoding performance as well as overall information recovery without penalizing the decoding of low-priority data, assuming high-priority data is no more than half of a message block. The cost is in the additional complexity required in the encoder. If extra computation resource is available at the transmitter, image, voice, and video transmission quality in terrestrial and space communications can benefit from accurate use of redundancy in protecting data with varying priorities.

Woo, Simon S.↗

Interferometer real time control development for SIM

This paper provides an overview of the architecture, design, integration, and test of the SIM flight interferometer real time control to meet challenging flight system requirements for the high processor throughput, low-latency interconnect, and precise synchronization to support microarcsecond-level astrometric measurements for greater than five years at 1 AU in Earth-trailing orbit.

real↗

A Multi-Satellite Framework to Rapidly Evaluate Extreme Biosphere Cascades: The Western US 2021 Drought and Heatwave

The increasing frequency and intensity of climate extremes and complex ecosystem responses motivate the need for integrated observational studies at low-latency to determine biosphere responses and carbon-climate feedbacks. Here, we develop a satellite-based rapid attribution workflow and demonstrate its use at a 1–2-month latency to attribute drivers of the carbon cycle feedbacks during the 2020-2021 Western US drought and heatwave. In the first half of 2021, concurrent negative photosynthesis anomalies and large positive column CO 2 anomalies were detected with satellites. Using a simple atmospheric mass balance approach, we estimate a surface carbon efflux anomaly of 132 TgC in June 2021, a magnitude corroborated 28 independently with a dynamic global vegetation model. Integrated satellite observations of hydrologic processes, representing the soil-plant-atmosphere continuum (SPAC), show that these surface carbon flux anomalies are largely due to substantial reductions in photosynthesis because of a spatially widespread moisture-deficit propagation through the SPAC between 2020 and 2021. A causal model indicates deep soil moisture stores partially drove photosynthesis, maintaining its values in 2020 and driving its declines throughout 2021. The causal model also suggests legacy effects may have amplified photosynthesis deficits in 2021 beyond the direct effects of environmental forcing. The integrated, observation framework presented here provides a valuable first assessment of a biosphere extreme response and an independent testbed for improving drought propagation and mechanisms in models. The rapid identification of extreme carbon anomalies and hotspots can also aid mitigation and adaptation decisions.

Causal model↗

Air Traffic Management TestBed: Messaging Performance

The Air Traffic Management (ATM) TestBed is an air traffic management modeling and simulation platform and framework developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The communication middleware, implemented in the TestBed framework layer, is a core feature for data message exchange. The feature provides an abstraction layer called Messaging Support to allow switching one middleware to another without a need to rebuild the simulation components. Messaging performance such as latencies, run durations, and throughputs are important factors. Low latencies can produce accurate results in high-fidelity and visualization models. Short run durations are preferred because better run efficiency can be achieved. High throughputs allow more runs to be executed concurrently. This technical memorandum studies and compares the messaging performance by running a full-day, fast-time simulation using three communication middleware as well as tweaking the default communication middleware settings used by the TestBed. Results indicate that the messaging performance could be improved by disabling either compression or persistence settings, while the run duration and throughput could be further improved by disabling both settings with a tradeoff of the message latencies increased by a factor of ten.

Chok Fung Lai↗

Short-Block Protograph-Based LDPC Codes

Short-block low-density parity-check (LDPC) codes of a special type are intended to be especially well suited for potential applications that include transmission of command and control data, cellular telephony, data communications in wireless local area networks, and satellite data communications. [In general, LDPC codes belong to a class of error-correcting codes suitable for use in a variety of wireless data-communication systems that include noisy channels.] The codes of the present special type exhibit low error floors, low bit and frame error rates, and low latency (in comparison with related prior codes). These codes also achieve low maximum rate of undetected errors over all signal-to-noise ratios, without requiring the use of cyclic redundancy checks, which would significantly increase the overhead for short blocks. These codes have protograph representations; this is advantageous in that, for reasons that exceed the scope of this article, the applicability of protograph representations makes it possible to design highspeed iterative decoders that utilize belief- propagation algorithms.

Divsalar, Dariush↗

PERSIANN-Unet: A Global Deep Learning Framework for Near-Real-Time Precipitation Estimation Using Infrared Data

Access to high-quality, high-resolution, near-real-time precipitation data is essential for hydrological and meteorological research and disaster mitigation. Traditional tools such as rain gauges and radar networks, though effective, have limitations, including sparse coverage in remote areas and high operational costs. Satellite data, with its global coverage and high spatial and temporal resolutions, mitigates limitations in coverage. Satellite precipitation products like Hydro Estimator (HE), Integrated Multi-satellitE Retrievals for Global Precipitation Measurement (IMERG), and Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks (PERSIANN) utilize both geosynchronous thermal infrared (IR) and passive microwave (PMW) data in their operation. PMW sensors offer detailed atmospheric profiles but suffer from higher latency, whereas IR sensors provide lower latency but only capture cloud-top information. Despite this constraint, IR data remains attractive for low-latency precipitation estimation. Recent advances in deep learning, particularly convolutional neural networks (CNNs), have further improved satellite precipitation retrievals. This study introduces PERSIANN-Unet (PUnet or PERSIANN V3), a quasi-global algorithm covering 60°N–60°S that combines IR data, monthly climatology, and the UNet architecture to produce half-hourly precipitation estimates at 0.04° resolution. The product is evaluated against HE, IMERG, and PDIR-Now for 2022–2023. Results show that PUnet closely matches its training target, IMERG V07 Final, at the global scale, and performance is further evaluated against Stage IV as a reference over CONUS. Training PUnet on IMERG (2016–2021) leverages a high-quality, integrated PMW IR-gauge precipitation product while developing an IR-based framework not reliant on PMW availability. By operating on a single global image, PUnet avoids tile partitioning and blending steps, reducing edge discontinuities, and produces more spatially consistent precipitation fields across hemispheres.

Phu Nguyen↗

Challenges Using the Linux Network Stack for Real-Time Communication

Starting in the early 2000s, human-in-the-loop (HITL) simulation groups at NASA and the Air Force Research Lab began using the Linux network stack for some real-time communication. More recently, SpaceX has adopted Ethernet as the primary bus technology for its Falcon launch vehicles and Dragon capsules. As the Linux network stack makes its way from ground facilities to flight critical systems, it is necessary to recognize that the network stack is optimized for communication over the open Internet, which cannot provide latency guarantees. The Internet protocols and their implementation in the Linux network stack contain numerous design decisions that favor throughput over determinism and latency. These decisions often require workarounds in the application or customization of the stack to maintain a high probability of low latency on closed networks, especially if the network must be fault tolerant to single event upsets.

Madden, Michael M.↗

The Impact of Traffic Prioritization on Deep Space Network Mission Traffic

A select number of missions supported by NASA's Deep Space Network (DSN) are demanding very high data rates. For example, the Kepler Mission was launched March 7, 2009 and at that time required the highest data rate of any NASA mission, with maximum rates of 4.33 Mb/s being provided via Ka band downlinks. The James Webb Space Telescope will require a maximum 28 Mb/s science downlink data rate also using Ka band links; as of this writing the launch is scheduled for a June 2014 launch. The Lunar Reconnaissance Orbiter, launched June 18, 2009, has demonstrated data rates at 100 Mb/s at lunar-Earth distances using NASA's Near Earth Network (NEN) and K-band. As further advances are made in high data rate space telecommunications, particularly with emerging optical systems, it is expected that large surges in demand on the supporting ground systems will ensue. A performance analysis of the impact of high variance in demand has been conducted using our Multi-mission Advanced Communications Hybrid Environment for Test and Evaluation (MACHETE) simulation tool. A comparison is made regarding the incorporation of Quality of Service (QoS) mechanisms and the resulting ground-to-ground Wide Area Network (WAN) bandwidth necessary to meet latency requirements across different user missions. It is shown that substantial reduction in WAN bandwidth may be realized through QoS techniques when low data rate users with low-latency needs are mixed with high data rate users having delay-tolerant traffic.

prioritization↗

The Impact of Traffic Prioritization on Deep Space Network Mission Traffic

A select number of missions supported by NASA's Deep Space Network (DSN) are demanding very high data rates. For example, the Kepler Mission was launched March 7, 2009 and at that time required the highest data rate of any NASA mission, with maximum rates of 4.33 Mb/s being provided via Ka band downlinks. The James Webb Space Telescope will require a maximum 28 Mb/s science downlink data rate also using Ka band links; as of this writing the launch is scheduled for a June 2014 launch. The Lunar Reconnaissance Orbiter, launched June 18, 2009, has demonstrated data rates at 100 Mb/s at lunar-Earth distances using NASA's Near Earth Network (NEN) and K-band. As further advances are made in high data rate space telecommunications, particularly with emerging optical systems, it is expected that large surges in demand on the supporting ground systems will ensue. A performance analysis of the impact of high variance in demand has been conducted using our Multi-mission Advanced Communications Hybrid Environment for Test and Evaluation (MACHETE) simulation tool. A comparison is made regarding the incorporation of Quality of Service (QoS) mechanisms and the resulting ground-to-ground Wide Area Network (WAN) bandwidth necessary to meet latency requirements across different user missions. It is shown that substantial reduction in WAN bandwidth may be realized through QoS techniques when low data rate users with low-latency needs are mixed with high data rate users having delay-tolerant traffic.

quality of service↗

Dual Purpose Simulation: New Data Link Test and Comparison With VDL-2

While the results of this paper are similar to those of previous research, in this paper technical difficulties present there are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL- 2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PCSMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA outperforms traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. We are testing a new and better data link that could replace CSMA with relative ease. Work is underway to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

Dual Purpose Simulation: New Data Link Test and Performance Limit Testing of Currently Deployed Data Link

While the results of this paper are similar to those of [I], in this paper technical difficulties present in [I] are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air-traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft-taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PCSMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA outperforms traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. So we have the tools to determine the traffic-loading conditions where traditional CSMA will fail, and we are testing a new and better data link that could replace it with relative ease. Work is currently being done to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

Dual Purpose Simulation: New Data Link Test and Comparison with VDL-2

While the results of this paper are similar to those of previous research, in this paper technical difficulties present there are eliminated, producing better results, enabling one to more readily see the benefits of Prioritized CSMA (PCSMA). A new analysis section also helps to generalize this research so that it is not limited to exploration of the new concept of PCSMA. Commercially available network simulation software, OPNET version 7.0, simulations are presented involving an important application of the Aeronautical Telecommunications Network (A TN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. The simulation involves 111 Air Traffic Control (ATC) ground stations, 32 airports distributed throughout the U.S., which are either sources or destinations for the air traffic landing or departing from the three cities, and also 1,235 equally equipped aircraft- taking off, flying realistic free- flight trajectories, and landing in a 24-hr period. Collision-less PCSMA is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, PCSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of PC SMA for implementing low latency, high throughput and efficient connectivity. Moreover, since PCSMA out performs traditional CSMA, by simulating with it, we can determine the limits of performance beyond which traditional CSMA may not pass. We are testing a new and better data link that could replace CSMA with relative ease. Work is underway to drastically expand the number of flights to make the simulation more representative of the National Aerospace System.

Robinson, Daryl C.↗

CSMA Versus Prioritized CSMA for Air-Traffic-Control Improvement

OPNET version 7.0 simulations are presented involving an important application of the Aeronautical Telecommunications Network (ATN), Controller Pilot Data Link Communications (CPDLC) over the Very High Frequency Data Link, Mode 2 (VDL-2). Communication is modeled for essentially all incoming and outgoing nonstop air-traffic for just three United States cities: Cleveland, Cincinnati, and Detroit. There are 32 airports in the simulation, 29 of which are either sources or destinations for the air-traffic of the aforementioned three airports. The simulation involves 111 Air Traffic Control (ATC) ground stations, and 1,235 equally equipped aircraft-taking off, flying realistic free-flight trajectories, and landing in a 24-hr period. Collisionless, Prioritized Carrier Sense Multiple Access (CSMA) is successfully tested and compared with the traditional CSMA typically associated with VDL-2. The performance measures include latency, throughput, and packet loss. As expected, Prioritized CSMA is much quicker and more efficient than traditional CSMA. These simulation results show the potency of Prioritized CSMA for implementing low latency, high throughput, and efficient connectivity.

Robinson, Daryl C.↗

HTMT-class Latency Tolerant Parallel Architecture for Petaflops Scale Computation

Computational Aero Sciences and other numeric intensive computation disciplines demand computing throughputs substantially greater than the Teraflops scale systems only now becoming available. The related fields of fluids, structures, thermal, combustion, and dynamic controls are among the interdisciplinary areas that in combination with sufficient resolution and advanced adaptive techniques may force performance requirements towards Petaflops. This will be especially true for compute intensive models such as Navier-Stokes are or when such system models are only part of a larger design optimization computation involving many design points. Yet recent experience with conventional MPP configurations comprising commodity processing and memory components has shown that larger scale frequently results in higher programming difficulty and lower system efficiency. While important advances in system software and algorithms techniques have had some impact on efficiency and programmability for certain classes of problems, in general it is unlikely that software alone will resolve the challenges to higher scalability. As in the past, future generations of high-end computers may require a combination of hardware architecture and system software advances to enable efficient operation at a Petaflops level. The NASA led HTMT project has engaged the talents of a broad interdisciplinary team to develop a new strategy in high-end system architecture to deliver petaflops scale computing in the 2004/5 timeframe. The Hybrid-Technology, MultiThreaded parallel computer architecture incorporates several advanced technologies in combination with an innovative dynamic adaptive scheduling mechanism to provide unprecedented performance and efficiency within practical constraints of cost, complexity, and power consumption. The emerging superconductor Rapid Single Flux Quantum electronics can operate at 100 GHz (the record is 770 GHz) and one percent of the power required by convention semiconductor logic. Wave Division Multiplexing optical communications can approach a peak per fiber bandwidth of 1 Tbps and the new Data Vortex network topology employing this technology can connect tens of thousands of ports providing a bi-section bandwidth on the order of a Petabyte per second with latencies well below 100 nanoseconds, even under heavy loads. Processor-in-Memory (PIM) technology combines logic and memory on the same chip exposing the internal bandwidth of the memory row buffers at low latency. And holographic storage photorefractive storage technologies provide high-density memory with access a thousand times faster than conventional disk technologies. Together these technologies enable a new class of shared memory system architecture with a peak performance in the range of a Petaflops but size and power requirements comparable to today's largest Teraflops scale systems. To achieve high-sustained performance, HTMT combines an advanced multithreading processor architecture with a memory-driven coarse-grained latency management strategy called "percolation", yielding high efficiency while reducing the much of the parallel programming burden. This paper will present the basic system architecture characteristics made possible through this series of advanced technologies and then give a detailed description of the new percolation approach to runtime latency management.

Sterling, Thomas↗