Performance Comparison of CUDA and OpenACC Based on Optimizations

Xuechao Li, Po-Chou Shih · 2018

Based on various optimizations, this paper presents a performance comparison between CUDA and OpenACC using 19 kernels in 10 benchmarks. The performance analysis focuses on programming models, optimization technologies and underlying compilers. It measures and compares kernel execution times and data transfer times to/from the GPU. In addition, it utilizes a Performance Ratio metric to conduct an objective comparison. The experimental results show that in general the PGI compiler is able to translate OpenACC kernels into object code that is slightly slower than hand-written CUDA codes for benchmarks that solve the same problem. Also, the data transfer time in OpenACC programs tends to be much faster than in CUDA, while the number of memcpy calls tends to be higher than in CUDA. Overall conclusions were found that OpenACC is a very reliable programming model and a good alternative to CUDA for accelerator devices. For the programs in our test corpus, OpenACC performs as well as CUDA, and in general, OpenACC is better for novices and for programmers targeting multiple platforms.

Read the paper · More papers on PaperTik