Leveraging Interaction Between Memory Footprint and Parallelism Degree for efficient GPU Portings

Mickaël Boichot, Adrien Roussel, Élisabeth Brunet, Patrick Carribault · 2025

Porting large HPC applications entirely on a Graphics Processing Unit (GPU) can be challenging. Some code portions are indeed unsuitable for GPU porting. Therefore, selecting the most profitable code parts for GPU porting is crucial. Many profiling tools address this issue, but their overhead is non-negligible for large HPC test cases. Moreover, the extracted code parts might not be the best candidates for any input set size. We present an approach that extrapolates the behavior of pre-selected code parts from different input set size runs on a target GPU. This enables developers to evaluate the application’s parallelism potential and memory footprint prior to GPU porting. We applied our approach to several HPC mini-applications and evaluated the extrapolations through a comparison to the existing GPU versions, as ground truth, on different vendors’ GPUs. Our results provide input set sizes of magnitude leading to GPU memory saturation and recommend which pre-selected code parts should be further studied for GPU porting.

Read the paper · More papers on PaperTik