Optimization of GEMV on Intel AVX Processor
Jun Liang, Yunquan Zhang · International Journal of Database Theory and Application · 2016
To improve the performance of BLAS 2 GEMV subroutine under the latest instruction set, Intel AVX, this paper presents a new approach to analyze the new generation instruction set and enhance the efficiency of current data-oriented math subroutines.The whole optimizing process involves memory access optimization, SIMD optimization and parallel optimization.Also, this paper shows the comparison between the traditional SSE instruction set and the AVX instruction set.Experiments show that the optimized GEMV function has obtained considerable increase on performance.Compared with the Intel MKL, GotoBLAS, ATLAS, this optimized GEMV exceeds these BLAS implementations from 5% to 10%.