DOE OSTI · 1867362
MemHC: An Optimized GPU Memory Management Framework for Accelerating Many-body Correlation
Abstract
The many-body correlation function is a fundamental computation kernel in modern physics computing applications, e.g., Hadron Contractions in Lattice quantum chromodynamics (QCD). This kernel is both computation and memory intensive, involving a series of tensor contractions, and thus usually runs on accelerators like GPUs. Existing optimizations on many-body correlation mainly focus on individual tensor contractions (e.g., cuBLAS libraries and others). In contrast, this work discovers a new optimization dimension for many-body correlation by exploring the optimization opportunities among tensor contractions. More specifically, it targets general GPU architectures (both NVIDIA and AMD) and optimizes many-body correlation’s memory management by exploiting a set of memory allocation and communication redundancy elimination opportunities: first, GPU memory allocation redundancy: the intermediate output frequently occurs as input in the subsequent calculations; second, CPU-GPU communication redundancy: although all tensors are allocated on both CPU and GPU, many of them are used (and reused) on the GPU side only, and thus, many CPU/GPU communications (like that in existing Unified Memory designs) are unnecessary; third, GPU oversubscription: limited GPU memory size causes oversubscription issues, and existing memory management usually results in near-reuse data eviction, thus incurring extra CPU/GPU memory communications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wang, Qihan, Peng, Zhen, Ren, Bin, Chen, Jie, Edwards, Robert G.. 2022-03-24. MemHC: An Optimized GPU Memory Management Framework for Accelerating Many-body Correlation. https://doi.org/10.1145/3506705
Cite the original work for its findings. Save a collection to share your selection of sources.