LBBGEMM: A Load-balanced Batch GEMM Framework on ARM CPU s
Cunyang Wei, Haipeng Jia, Yunquan Zhang, Kun Li, Luhan Wang · 2022
The trend in modern high performance computing is to decompose a large linear algebraic problem into many small problems that can be solved independently. Although there are many studies to obtain near-peak performance for large-scale dense matrix operations, it is not sufficient for batch operations with small matrices. The study of small GEMM kernel optimization and load balanced scheduling of batch operations on ARM processors is not enough. In this paper, we present LBBGEMM, a load-balanced batch GEMM framework for optimizing large groups of variable-size small GEMM to boost near-optimal performance based on ARMv8 architecture. The LBBGEMM is divided into the install-time stage and the run-time stage. At the install-time stage, we analyze the characteristics of each transposition mode to design the high-performance small GEMM kernel without data packing. This strategy greatly reduces the memory access overhead. In addition, we optimize instruction scheduling and instruction selection carefully to achieve optimal performance. The run-time stage provides a comprehensive auto-tuning process for batch GEMM by using a tiling designer and a pre-grouped dynamic scheduling algorithm. The tiling designer generates high-performance execution plans for each group of matrices with different sizes. Then we divide the large group of GEMM operations into small task groups. These task groups are assigned to threads for execution in the form of command queues through our proposed dynamic mapping between threads and tasks. This pre-grouped multi-thread task scheduling algorithm greatly improves the speedup of multi-thread. The experiments show that LBBGEMM could achieve significant performance improvements in batch GEMM compared with other mainstream BLAS libraries.