Engineering PapersSearch

SEARCH · Engineering Papers

Results for “POINT-TO-POINT COMMUNICATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)

Phloem

Phloem is a Message Passing Interface (MPI) micro-benchmarking suite featuring sub-communicator collectives, methods for finding slow links on MPI interconnects, and point-to-point MPI benchmarks, including a messaging rate benchmark. All of the benchmarks except for ones related exclusively to finding slow links are GPU-aware via the Umpire resource management library.

Moody, AdamT [Lawrence Livermore National Laborato

7-8 GHz Point-to-Point Testing

Wireless spectrum is a limiting resource for continued economic growth in the United States. As discussed in the recent National Spectrum Strategy (NSS), wireless spectrum underpins several aspects of the U.S. economy and the demand for additional spectrum is driving the need for realizing spectrum sharing to enable continued development. At the same time, wireless spectrum is an essential foundation of critical energy infrastructure, including electric, oil, and natural gas resources. In particular, wireless point-to-point (P2P) links are the backbone of vast infrastructure networks that enable the flow of sensor and control information needed to manage critical energy sector infrastructure in the United States. These links will only become more important as the energy sector incorporates more diverse sensing and more efficient control mechanisms, which increases system complexity and data volume requiring more resilient communications. Therefore, the continued economic development of the U.S. depends on determining novel approaches to spectrum management that balance both broad access for advanced wireless technologies and resilience for critical infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION

A multigigabit link layer protocol for single picosecond latency determinism using AMD ultrascale+ GTH and GTY transreceivers

recision timing distribution systems require determin istic and repeatable high speed serial data link latency. The EIC Timing Data Link will employ a specialized link layer protocol for deterministic multigigabit communication us ing AMD Ultrascale+ GTH and GTY transceivers. Link la tency must also be measurable to accurately compensate in real time for latency variations of the physical medium. The point-to-point link is full duplex and fully synchronous with a latency control algorithm to align the internal clocks of each system with picosecond resolution. Deterministic clock domain crossing between systems is ensured using static timing analysis of the elastic buffer control signals.

43 PARTICLE ACCELERATORS

Digital Grid Twin–Direct Communication Scheme Test Bed for Assessing Relay-to-Relay Radio Antenna and Optical Fiber Performance and Misoperations

This study introduces a novel “Digital Grid Twin–Direct Communication Scheme” test bed. This advanced platform evaluates point-to-point communication between transmitter and receiver relays with optical fiber and radio omnidirectional antenna systems, implemented at the Advanced Protection lab in the Grid Research Innovation and Development Center at Oak Ridge National Laboratory. The increased diversity of energy sources has led to more protective relay misoperations. In North America, microgrid protection schemes now use point-to-point communication along distribution lines between relays to implement advanced logic in nonradial grids that include both high- and low-inertia generators. This trend challenges utilities to minimize misoperations while ensuring rapid fault clearance and accurate selectivity coordination between primary and backup relays. This study assesses relay-to-relay communication schemes by introducing an advanced testing platform based on a digital grid twin protection test bed using a synchronized time source system. The platform evaluates the communication system using radio antennas or optical fiber links by integrating protective relays that operate breakers within the digital twin and record relay events and communication signals. In the experiments, transmitter and receiver relays were configured with inverse time overcurrent and breaker trip detection logic to assess the total time of the communication protection schemes based on the sum of the relay protection element operating time, radio latency, propagation delay, baud rate delay, and relay processing time. These delays were derived from recorded relay events and communication signals from the interface of a real-time simulator set as a digital grid twin. The test bed successfully simulated various electrical faults along a distribution line while ensuring effective and reliable point-to-point communication between transmitter and receiver relays. The radio antenna communication system exhibited latency because of the radio. This latency depends on the baud rate setting and type of radio application; in general, the higher the baud rate, the lower the radio latency. The measured radio latency (for Mirrored Bits with an encryption card at 9,600 bps) was about 9–10 ms. Additionally, calculated propagation delay per mile for radio antennas and optical fiber was 5.36 µs/mi and 8.04 µs/mi, respectively. Optical fiber communication did not demonstrate radio latency. Instead, the protection element operating time depends mainly on the protection logic function set in the relay, and the relay processing time depends on the processing rate of the relay in samples per power system cycle.

24 POWER TRANSMISSION AND DISTRIBUTION

Digital Grid Twin–Direct Communication Scheme Test Bed for Assessing Relay-to-Relay Radio Antenna and Optical Fiber Performance and Misoperations

This study introduces a novel “Digital Grid Twin–Direct Communication Scheme” test bed. This advanced platform evaluates point-to-point communication between transmitter and receiver relays with optical fiber and radio omnidirectional antenna systems, implemented at the Advanced Protection lab in the Grid Research Innovation and Development Center at Oak Ridge National Laboratory.

99 GENERAL AND MISCELLANEOUS

Efficiently Computable Limits on EPR Pair Generation in Quantum Broadcast Channels

We investigate the generation of EPR pairs between three observers in a general causally structured setting, where communication occurs via a noisy quantum broadcast channel. The most general quantum codes for this setup take the form of tripartite quantum channels. Since the receivers are constrained by causal ordering, additional temporal relationships naturally emerge between the parties. These causal constraints enforce intrinsic no-signalling conditions on any tripartite operation, ensuring that it constitutes a physically realizable quantum code for a quantum broadcast channel. We analyze these constraints and, more broadly, characterize the most general quantum codes for communication over such channels. We examine the capabilities of codes that are fully no-signalling among the three parties, positive partial transpose (PPT)-preserving, or both, and derive simple semidefinite programs to compute the achievable entanglement fidelity. We then establish a hierarchy of semidefinite programming converse bounds -- both weak and strong -- for the capacity of quantum broadcast channels for EPR pair generation, in both one-shot and asymptotic regimes. Notably, in the special case of a point-to-point channel, our strong converse bound recovers and strengthens existing results. Finally, we demonstrate how the PPT-preserving codes we develop can be leveraged to construct PPT-preserving entanglement combing schemes, and vice versa.

FOS: Physical sciences