HWSA: A High-Ratio Weight Sparse Accelerator for Efficient CNN Inference
Xuejing Dai, Jinze Zhang, Zhongfeng Wang, Jun Lin · IEEE Transactions on Circuits and Systems I Regular Papers · 2025
Pruning has emerged as an effective technique for compressing convolutional neural networks (CNNs) by eliminating redundant weights, achieving lightweight models with negligible loss in inference accuracy. To leverage the sparsity for acceleration, many accelerators built for sparse CNNs have been developed. Existing hardware accelerators can perform well with structured or customized sparsity patterns. However, when facing the unstructured sparsity which can achieve higher compression rates, the corresponding hardware always suffers from insufficient utilization of computational resources, severe load balance between process elements, and significant overhead in logic resources, which results in a reduction of throughput, making it difficult to leverage high ratio sparsity for acceleration effectively. An efficient CNN inference accelerator that can handle both structured and unstructured sparse networks is proposed to address these issues. By flexibly employing multiple parallel computation methods combined with carefully developed sorting algorithms, the proposed architecture mitigates the hardware utilization inefficiencies caused by unstructured sparsity. Through a hardware-software methodology, new sparsity rules are introduced, nearly eliminating the load imbalance issue. The proposed Processing Element (PE) architecture can effectively select inputs for sparse networks while reducing the overhead of logic resources. The proposed architecture is implemented on the XCVU9P FPGA, achieving a frequency of 200 MHz. It achieves a computational throughput of 350.49 GOPs and 326.01 GOPs on ResNet-50 and ResNet-152, respectively, demonstrating a 1.2-$2.1\times $DSP efficiency improvement compared to previous works.