An Efficient Batch GEMM Framework for Deep Neural Networks
Siliang Suo, Zhihong Liang, Xiaowei Zhao · 2025
GEneral Matrix Multiplication (GEMM) is a fundamental mathematical operation for various applications, including high-performance computing and machine learning. Within deep neural networks, batched, irregular, and small-scale matrix multiplication operations can consume a significant amount of computational time. The traditional rocBLAS library is typically optimized for large matrices. Its computing pattern exhibits low utilization of computing units and poor memory access efficiency when handling small-size matrices, and it lacks the ability to adapt dynamically. To handle the above problems, we propose an efficient batch GEMM framework, including prefetching and double buffering design, various-size tile kernels, and the tiling algorithm. The experimental results on MI210 show that the proposed framework achieves a performance improvement of $1.15 \times$ (up to $1.41 \times$) over rocBLAS.