GPUPerfML: A Performance Analytical Model Based on Decision Tree for GPU Architectures

Ran Zheng, Qingyue Hu, Hai Jin · 2018

GPU has been applied to many fields for its excellent parallel computing ability, but the complexity of GPU architecture makes it hard to find the performance bottlenecks of GPU applications. Although profiling tools exist, they only provide a large number of data that programmers can not well understand. Some mathematical modelling methods are presented for analysis, but they are often too complicated and time-consuming for users. A performance evaluation method named GPUPerfML is proposed, which combines decision tree and theoretical analytical model to locate performance bottlenecks of GPU applications and guide application optimization. Based on its feature selection, decision tree is used to extract the sorted influential features from a large number of features. The proposed theoretical analysis model defines the categories of performance issues and establishes mapping relationships between features and performance issues, so as to quickly and directly identify performance bottlenecks. Besides, theoretical analysis can be used to guide decision tree building and features analysis, which improves and guarantees the accuracy of the proposed method. The experiments on three common applications (Matrix transpose, Parallel Reduction, and BFS) show that it can identify performance bottlenecks accurately and quickly, to help programmers write high-efficient programs conveniently.

Read the paper · More papers on PaperTik