HMCSP: Reducing Transaction Latency of CSR-based SPMV in Hybrid Memory Cube
Cheng Qian, Bruce R. Childers, Libo Huang, Qi Yu, Zhiying Wang · 2018
Sparse Matrix Multiplication Vector (SPMV) plays a significant role in sparse linear algebra. Based on the high parallelization of matrix multiplication, SPMV has been accelerated with GPUs, Intel MIC, and FPGAs. The Micron Hybrid Memory Cube (HMC) is a highly parallel device that has atomic operations which support processing in memory (PIM). In this paper, we propose HMCSP, which extends the HMC's existing PIM capability to reduce the memory transaction latency of SPMV. By taking advantage of atomic operations and data prefetch, HMCSP reduces memory transaction latency of SPMV by 49.7% compared to a conventional HMC.