SCRA: Systolic-Friendly DNN Compression and Reconfigurable Accelerator Co-Design

Xun Zhang, Chengliang Wang, Xingquan Piao, Ao Ren, Zhetong Huang · 2023

Pruning has become an extremely powerful and effective technique to compress and accelerate sophisticated deep neural networks on resource-constrained platforms. Existing pruning and quantization methods aim to reduce parameter size and computation while maintaining accuracy. However, it is impractical to ignore hardware implementation when focusing on the accuracy of DNNs through pruning and quantization strategies. Moreover, a reduction in computational size does not necessarily result in an increase in inference efficiency.In order to strike a balance between DNN accuracy, inference efficiency, and hardware implementation, we propose SCRA, an efficient algorithm-architecture co-design that leverages stripe-based structured pruning strategy to reduce the data movement and fit the data flow of systolic arrays (SAs). Besides, considering the characteristics of the DSP48E2 in the FPGA, we present a quantization strategy for parallel computation of mixed-precision data, achieving acceleration of mixed-precision DNNs computation. Furthermore, we put forward an accelerator to fully exploit the benefits of the pruning strategy and parallel computation. Experimental results demonstrate that the proposed pruning strategy can achieve higher pruning rates without significant loss in model accuracy, and our accelerator achieves 4.79x speed up compared to the baseline.

Read the paper · More papers on PaperTik