Predicting Execution Time of CUDA Kernel Using Static Analysis
Gargi Alavani, Kajal Varma, Santonu Sarkar · 2018
With the growing demand for performance-oriented problems, programmers routinely execute the embarrassing parallel part of the application (GPU kernels) in a GPU in order to achieve signi?cant speedup. These applications are becoming complex and long-running which makes it energy inef?cient. Anticipating its execution time can help the developers to ?x the inef?cient code before running it. In this paper, we propose an approach to predict the execution time of a GPU kernel without the need of executing it. We build an analytical model to predict the execution time of a GPU kernel by analyzing the intermediate PTX code of a CUDA kernel. Our experimental analysis of a set of benchmarks shows that for 45 applications the estimated execution time has the mean absolute error of 26.86% when compared to the actual execution time. Mean absolute error for benchmarks belonging to Dynamic programming dwarf is minimum, followed by Dense Linear Algebra benchmarks.