A bandwidth in-sensitive low stall sparse matrix vector multiplication architecture on reconfigurable FPGA platform

Syed Mohsin Ali, Wang Shaojun, Ning Ma, Yu Peng · 2017

Sparse Matrix with dense Vector Multiplication (SpMxV) is a key computational kernel for most mathematical high performance computing problems. SpMxV implementation on FPGAs provides better performance but memory bandwidth limitations and pipeline stalling are still ongoing research concerns. This research proposes an FPGA optimized single bit stream column major sparse compression format that ensures decreased memory bandwidth and storage requirement. The architecture ensures low pipeline stalling by effectively performing 3 floating point operations per cycle per processing element. The proposed architecture is demonstrated with known benchmarks on ZYNQ-ZC702 FPGA. Result shows that on average our approach achieves compression ratios of 3.0 to 5.3 with respect to conventional CSR format. And peak compression ratio of 20.0 can be achieved for certain matrices used in our research. Low stall architecture proves 40 to 150% depreciation in stall cycles providing better overall efficiency.

Read the paper · More papers on PaperTik