A High-performance SpMV Accelerator on HBM-equipped FPGAs

Tao Li, Li Shen, Shangshang Yao · 2022

Sparse Matrix-Vector Multiplication (SpMV) is an important kernel that is widely used in science and engineering applications. The features of SpMV, such as high memory-intensiveness and many different access patterns, cause that the performance of SpMV to be bounded by the limited bandwidth between memory and processing units (PEs). High-bandwidth memory (HBM) is a novel memory system to overcome the bandwidth bottleneck. In this paper, we propose a HBM-based SpMV accelerator. To make full use of the high bandwidth provided by HBM, we design a highly parallel PE array and implement a high-frequency pipeline inside the PE. For each PE, we integrate L1 cache to exploit the data locality in vector. We propose two data layout strategies, namely a row merging algorithm to exploit the inter-row data locality and a row assignment algorithm to achieve workload balance among PEs. Our design is implemented using a FPGA card with 8GB HBM2 memory. Compared to the baseline CPU SpMV implementation, our accelerator can obtain a 6.39x performance speedup on average.

Read the paper · More papers on PaperTik