DOE OSTI · 1996148
GPU Profiling and Optimizing xRAGE (Final Report)
Abstract
Our project’s objective is to increase the efficiency of GPU-enabled kernels in xRAGE. To do so, we conduct GPU profiling with NSight Systems on xRAGE tests unsplit_sod_1d and unsplit_sedov_2d to identify bottlenecks and understand the behavior of the GPU during code execution. Next, we analyze these generated GPU profiles to locate the lines of code whose optimization have the most potential for improving runtime. We replicate the structure of the code in smaller test problems that are easier to understand, edit, and run quickly. Within these test problems, we implement two different methods of improving performance: transformation of nested loops into a single MDRangePolicy and hierarchical parallelization using teams of threads. Both methods show speedups in the test code, and after transferring them to xRAGE, they both show up to 30x speedups on various computing platforms. Profiling the edited versions of xRAGE reveals that the GPU successfully executed the bottlenecks with greater efficiency
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zeng, Jason Luo, Gonzalez, Ivan Eduardo. 2023-08-18. GPU Profiling and Optimizing xRAGE (Final Report). https://doi.org/10.2172/1996148
Cite the original work for its findings. Save a collection to share your selection of sources.