A Compiler Framework for Optimization of Affine Loop Nests for General Purpose Computations on GPUs

Muthu Manikandan Baskaran, Uday Kumar Reddy Bondhugula, Sriram Krishnamoorthy, Jagannathan Ramanujam, Atanas Rountev, Ponnuswamy Sadayappan · 2015

GPUs are a class of specialized parallel architectures with tremendous computational power. The new Compute Unified Device Architecture (CUDA) programming model from NVIDIA facilitates programming of general purpose applications on NVIDIA GPUs. However, there are various performance-influencing factors specific to GPU architectures that need to be accurately characterized to effectively utilize the parallel computing power of GPUs and improve the performance of applications on GPUs. Often these factors are tightly coupled, making their effective tuning a significant challenge. In addition, program-specific optimizations such as tiling, loop unrolling, etc. involve performance trade-offs on GPUs that are difficult to characterize accurately using performance models. In this paper, we develop an automatic compiler framework for generating efficient parallel programs on GPUs for given input regular programs. The framework generates program transformations (using the general polyhedral model) that enable efficient execution over GPUs, and employs a model-driven empirical optimization approach to find optimal values for system parameters that maximize performance, as well as the best tile sizes and loop unroll factors. Experimental results show a significant improvement in performance for kernels generated using our framework, and provide new insights into performance optimizations for GPUs.

Read the paper · More papers on PaperTik