Performance and Portability Analysis of OPS Structured Grid Parallel Library
Fengting Chen, Yonggang Che, Wenke Wang, Jian Ma · 2021 IEEE 3rd International Conference on Frontiers Technology of Information and Computer (ICFTIC) · 2021
The complexity and diversity of high-performance computer architectures have brought great challenges to parallel application development. Using DSL (Domain Specific Language) to achieve multi-platform automatic parallelization is a solution to this problem. This paper tests and analyzes the multi-platform parallel performance of OPS (Oxford Parallel Library for Structured Mesh Solvers), a typical DSL framework for scientific applications. Two representative structured mesh applications are used in our evaluation. The performance of their MPI, hybrid MPI/OpenMP and CUDA versions generated by OPS are evaluated on the Intel Xeon E5-2600 V3 CPU and the NVIDIA Tesla K80 GPU. The performance of OPS generated codes are compared against the corresponding manually parallelized and optimized codes. The results show that for the two applications, OPS based approach can generate parallel codes with comparable performance to manually parallelized and optimized codes. This paper further analyzes the differences between OPS-based multi-platform parallel code generation and manual parallelization approaches and their impact on performance. It also quantitatively evaluates the cross-platform portability of OPS, and points out some possible directions of improvement.