Auto-Tuning OmpSs-OpenCL Kernels Across GPU Machines
Vinoth Krishnan Elangovan, Rosa M. Badia, Eduard Ayguadé · 2015
Perceived benefits of the heterogeneous computing models have led to GPUs becoming an indispensable part of existing/future HPC infrastructures. The semiconductor industry has enabled a continuous increase in computing capabilities of GPUs across successive product generations. In this light, performance portability of applications across different generations of GPUs with minimal programmer intervention, is desired. There is, however a lack of tool-chain support that can help with porting applications across GPUs with myriad computing capabilities. In this paper, we propose an Auto-Tune tool for OmpSs-OpenCL kernels to ease the process of porting applications across different generations of GPUs. The proposed tool focuses on modifying the OpenCL kernel execution configuration based on GPU specifications and does not require any user intervention providing reasonable performance portability. We evaluate our proposal on 7 different benchmarks with varied characteristics ported across three different GPU architectures. In addition, the Auto-Tune tool integrated with OmpSs-OpenCL offered a maximum performance gain of 10% for Matmul workload at runtime.