Performance of the Modern Parallel Programming Approaches: A Case Study
Volodymyr Fedynyak, Oleksa Hryniv, Bohdan Vey, Oleg Farenyuk · 2023
This research is devoted to a quantitative comparison of the performance of several parallel programming approaches and compares their computational performance. Comparison is performed for the Computational Dynamics Problem solved by the MacCormack scheme. Parallel computation properties of this sample problem task are well-understood. The parallel programming techniques were chosen considering the recent trends in high-performance computing. Both high-level framework-based implementations (using OneAPI DPC++ and ArrayFire) and low-level implementations (based on CUDA C++) are reviewed and their performance is compared. Additionally, single SMP systems with multiple CUDA-capable GPUs were studied using GPUDirect and Unified Memory technologies. Wall-time was used as a performance metric. The comparison was performed using Student’s t-test for the Gauss-distributed experimental results and the non-parametric Wilcoxon signed-rank test – for other distributions. The results show that CUDA-based solutions outperform other approaches though development time considerations often can favor more high-level approaches.