A Multi-Tiered Autotuner for Portable Heterogeneous-Compute Interfaces
Shahriar Ahmed Zisan, Mohammad Nooruddin, Apan Qasem · 2025
Writing performance portable code has been a longstanding challenge in the high-performance computing (HPC) community. This challenge has been exacerbated by the inclusion of accelerators in HPC systems from multiple competing vendors. To address performance portability issues on accelerator-based heterogeneous systems, AMD and Intel have recently introduced suites of APIs capable of generating code for different accelerators with reasonable performance. While these APIs improve programmer productivity, achieving optimal performance still requires significant manual tuning effort.This paper proposes a multi-tiered approach to autotuning CUDA applications ported to AMD architectures via the Heterogeneous-Compute Interface (HIP). We begin by deconstructing the HIP compilation process, identifying key transformations in converting CUDA to HIP. We pinpoint the native compiler optimizations that most influence the performance of CUDA codes when executed on AMD architectures with the HIP runtime interface. We then construct a combined search space around these transformations and their parameters and explore it using a Genetic Algorithm. Our autotuner is the first system to investigate the HIP autotuning space and includes a regression-based strategy to quickly narrow down the search space to the most promising regions. Experimental results with the Rodinia benchmark suite demonstrate that our autotuning strategy achieves an average speedup of 1.7, with performance improvements of up to a factor of seven on certain applications.