DGEMM Optimization Oriented to ARM SVE Instruction Set Architecture
Wei Yi, Lin Deng, Sizheng Sun, Sisi Li, Li Shen · 2023
In current computer performance evaluation, linear mathematics library is an important test program and matrix multiplication is its main calculation part. Matrix multiplication algorithm contains multi-layer loops and can be parallelized flexibly. It is very suitable to run on multi-core processor with vector registers. With the popularity and application of ARMV8 vector processors in the field of high performance computing, the peak performance of processors is required to be higher. In this paper, we propose a double precision general matrix multiplication DGEMM vectorization method based on OpenBLAS and implemented on Phytium processor. The main work includes: the data block size is optimized for the storage structure of the processor and the data rearrangement is redesigned; based on the SVE vector instructions, a mathematical model is established to find the optimal kernel function size; an efficient assembly kernel is realized by using computing instructions to hide the memory delay. The experimental results show that the optimized code improves the measured performance of OpenBLAS original DGEMM algorithm from 45.07% of the theoretical peak performance of single core to 80.18%.