HR-SpMM: Adaptive Row Partitioning and Hybrid Kernel Design for Sparse Matrix Multiplication
Qi Wang, Y.P. Wang, Yi Luo, Rong Luo, Pingping Tang · 2025
Sparse Matrix-Matrix Multiplication (SpMM) plays a critical role in high-performance computing and applications like Graph Neural Networks (GNNs).However, due to the sparsity and irregularity of real-world data, optimizing SpMM performance on modern GPUs has remained a significant challenge.Existing methods often involve trade-offs between load balancing and hardware utilization, making it difficult to efficiently handle both long and short rows in sparse matrices.To address these issues, we propose HR-SpMM, a lightweight framework based on adaptive row partitioning and hybrid kernel design.HR-SpMM divides sparse matrix rows into two categories: long rows and short rows, leveraging Tensor Cores and CUDA Cores, respectively, to optimize computational efficiency.Long rows are further partitioned into fixed-size blocks to fully align with the hardware characteristics of Tensor Cores, while short rows adopt a flexible