Performance Optimization on GPGPU & Multicore CPU Using Roofline Model

Noor M. Allayla, Shefa A. Dawwd · IOP Conference Series Materials Science and Engineering · 2021

Abstract The roofline model introduced in this paper to evaluate the best optimized platform for training the neural network that used to recognize handwritten digits under multicore CPU and general-purpose GPU (GPGPU) as hardware environment. The pattern parallel training technique for MNIST dataset is applied. The parallel network training of MNIST using different data layout of multicore CPU and GPGPU is presented. Different bottlenecks have been explained by applying the roofline model. The most suitable platform is selected according to layouts and constrains either for memory or computation bounds. The computational intensity of all rooflines is moved toward right, then the performance is increased. As a result of optimization, and with the diversity of the available data size, core number, operational strength, the most suitable hardware platform is selected.

Read the paper · More papers on PaperTik