Efficiently Running SpMV on Multi-Core DSPs for Block Sparse Matrix
Deshun Bi, Xiaowen Tian, Shengguo Li, Dezun Dong · 2023
Sparse Matrix-Vector Multiplication (SpMV) is a fundamental operation in sparse computations. Although many techniques have been developed to speed up SpMV, optimizing this process on low-power multicore digital signal processors (DSPs) has often been neglected. This paper presents the FT-M7032, a cutting-edge CPU-DSP hybrid multi-core processor. We assess the data transfer efficiency among various units to identify performance constraints of SpMV on multicore DSPs. Based on our evaluation, we develop a method for block sparse matrices, namely SpMV_BLOCK, which can break the bandwidth bottleneck of SpMV and achieve significant performance gains. We then propose a load-balancing strategy for each thread by using binary search and devise a pipeline that overlaps data transfers and computations to improve SpMV performance. To measure our method’s effectiveness, we compared its performance against a baseline on the FT-M7032’s general-purpose CPU cores. Our experiments show that our approach delivers a notable 5.80 × speedup over the baseline.