Performance Estimation on Heterogeneous Systems: Making the most of Static Analysis

K Vanishree, Madhura Purnaprajna · 2020

Heterogeneous Computing System (HCS) comprising of accelerators such as GPU, FPGA and DSP are extensively used in the parallel computing domain. The diversity in their micro-architectures makes them suitable for the various parallel scientific applications. Most of the existing systems that address data distribution in HCS heavily depend on the target architecture that limits the design space exploration to a known device micro-architecture. In contrast, this work uses static code analysis to develop a target-independent performance model to suggest the suitability of a data-parallel regular application to CPU or GPU in an heterogeneous node. This model uses information available at compile time to estimate the performance and the objective is to statically obtain relative performance. With the performance estimates for both the CPU and GPU code for varied problem sizes, an application is classified as CPU-GPU or GPU-only. Furthermore, the approach also gives an optimal data distribution ratio for the application. This approach is evaluated using data-parallel applications that have varied speedups on GPU w.r.t. multi-core CPU. Using the proposed technique, an average performance improvement of 38.44% is seen across CPU-GPU benchmarks, with the co-execution of CPU+GPU as compared to CPU alone.

Read the paper · More papers on PaperTik