DOE OSTI · 3013469
Large-Message All-to-All Communication at Frontier Scale
Abstract
Near the full scale of exascale supercomputers, latency can dominate the cost of all-to-all communication even for very large message sizes. We describe GPU-aware all-to-all implementations designed to reduce latency for large message sizes at extreme scales, and we present their performance using 65536 tasks (8192 nodes) on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. Two implementations perform best for different ranges of message size, and all outperform the vendor-provided MPI_Alltoall. Our results show promising options for improving implementations of MPI_Alltoall_init.
Keep this discovery
Explore connections, maps & timelines
White, Trey [ORNL] (ORCID:000900052186075X). 2025-11-01. Large-Message All-to-All Communication at Frontier Scale. https://doi.org/10.1145/3731599.3767389
Cite the original work for its findings. Save a collection to share your selection of sources.