FullSparse: A Sparse-Aware GEMM Accelerator with Online Sparsity Prediction

Jiangnan Yu, Fan Yang, Hanfei Wang, Yuxuan Qiao, Zheng Yu Wu, Xiankui Xiong, Xiao jing Yao, Haidong Yao, Yecheng Zhang · 2024

Leveraging sparsity optimizes storage and computation for resource-constrained devices in Deep Learning Neural Networks (DNNs). While neural networks naturally incorporate sparsity through operations like ReLU and quantization, diverse sparsity levels (0.2% to 99%) pose challenges for the design of computational units. In this paper, we provide an energy-efficient GEMM accelerator named FullSparse which is designed for diverse applications, accommodating varying sparsity levels in matrix multiplication (0.2% to 99%). This paper introduces three features for nuanced sparsity support: multi-sparsity control, predictive result sparsity, and a multi-sparsity-compatible PE array. Experimental evaluations affirm that our implementation while ensuring adaptability to sparsity, exhibits superior computational power comparable to the existing designs.

Read the paper · More papers on PaperTik