Optimization and Performance Modeling of Stencil Computations on ARM Architectures
Kaifang Zhang, Huayou Su, Peng Zhang, Yong Dou · 2020
Stencil Computation has long been an omnipresent kernel of a wide range of scientific and engineering applications. There is much work investigating the stencil performance on x86 processors and accelerators such as GPU. Meanwhile, ARM processors for HPC have been highlighted more and more with Fugaku of Japan becoming the first number one system on the TOP-500 list. In this paper, we focus on modeling and optimizing the performance of stencil computation on ARM architectures. Specifically, we proposed a performance model for stencil computation based on tiling optimization to guide the optimal configuration of tiling parameters for good cache reuse. We validate the proposed model with the Performance Monitor Unit (PMU) provided by the ARM processor. Experimental results show that the prediction error of the execution time can be lower to 1.37%. Furthermore, we can achieve a maximum of 1.26 × speedup compared to the naive implementation according to the optimization parameters provided from the model.