A Framework for Performance Analysis and Optimization for GPU Kernel Programs using Linear Performance-Breakdown Model
Martell-Mario Alberto Chapa · Institutional Repositories DataBase (IRDB) · 2015
Graphics Processing Units (GPU) have evolved into computational devices with Teraflop performance capabilities.However, a programmer writing code for GPUs faces a mode challenging task than when programming for a CPU.First, it is necessary to correctly identify parallel computation.Second, Optimize the placement and movement of data in the multi-level memory hierarchy of the GPU.In addition, a program that have been optimized for one architecture might fail to reach the expected performance on a different GPU architecture.The described scenario leads to a situation where large amounts of time and effort are needed to produce a correct and efficient GPU program.GPU application programming can be aided by performance models and frameworks, and in this document we describe our performance modeling framework for GPU programs.We propose the Linear Performance-Breakdown Model (LBPM), a tool designed to help the programmer to locate performance bottlenecks, facilitating the optimization process and a framework to apply the LBPM model.The objective of the framework is to provide a performance tunning tool to facilitate the optimization of GPU kernel programs.By extracting the breakdown of the execution time of kernel programs into three main components (global memory transfers, local memory transfers and floating-point operations time) the model can serve as a tool to guide optimization efforts We demonstrate the effectiveness of our model to calculate the breakdown of performance by applying it to several case-tests: SGMM, FFT, Reduction.We confirmed the modeling methodology works with two different GPU devices: A8-3870 AMD Accelerated Processing Unit (GPU) and a GTX 660 Nvidia GPU.