Engineering PapersSearch

DOE OSTI · 3682548

Profile Generation for GPU Targets

Abstract

GPU accelerators are ubiquitous, but their ecosystem is far less evolved than the host one. Compiler heuristics are often tuned for CPUs and reused for GPU. Similarly, tooling and more evolved optimization techniques are historically not available on GPU targets. In this work, we address one of these shortcomings and enable profile generation and profile-guided optimizations (PGO) for GPU targets. While this is only a single step towards a CPU equivalent ecosystem for offload devices, it shows how old misconceptions on the limitations of GPUs are often not warranted anymore. Through our implementation in LLVM/Offload, we enable device-side PGO for full scientific applications and open up tooling opportunities, including code coverage analysis and compiler-built-in roofline analysis. Our evaluation highlights the performance implications of profile generation, the insights gained from these profiles, and the (missed) opportunities in utilizing the information for GPU compilation.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

McDonough, Ethan Luis [Lawrence Livermore National Laboratory (LLNL)] (ORCID:0009000471513433), Denny, Joel [ORNL] (ORCID:0000000196995599), Doerfert, Johannes [Lawrence Livermore National Laboratory (LLNL)] (ORCID:0000000178708963). 2026-09-01. Profile Generation for GPU Targets. https://doi.org/10.1007/978-3-032-06343-4_7

Cite the original work for its findings. Save a collection to share your selection of sources.