Exploiting Tensor Cores in Sparse Matrix-Multivector Multiplication via Block-Sparsity-Aware Clustering
Eunji Lee, Yoonsang Han, Gordon Euhyun Moon · 2024
Sparse Matrix-Multivector (SpMM) multiplication is a key kernel for deep learning models and scientific computing applications. However, achieving high performance for SpMM on GPUs is challenging because of the irregular distribution of non- zero elements and irregular memory accesses in sparse matrices. In this paper, we propose a novel sparse matrix reordering algorithm based on the block sparsity patterns in rows to improve data locality for SpMM. To accelerate our reordering algorithm, we develop a parallel implementation based on GPUs. The high-density tiles in the reordered matrix are used to perform accelerated dense matrix multiplication by leveraging Tensor Cores on GPUs. Experimental results on a large number of sparse matrices demonstrate that our TC-SpMM achieves an average speedup of 3.4x and a peak speedup of 20.77 x over cuSPARSE.