A novel hardware acceleration method for Deep Neural Networks
Dingjiang Yan, Fangzhou He, Fenglin Cai, Mingfei Jiang, Jie Li · 2023
The huge computational complexity and memory requirement of deep neural networks (DNNs) impede their deployability on resource-constrained edge-computing devices. To tackle this challenge, numerous model compression solutions have emerged, where Power-of-Two (PoT) quantization scheme demonstrates its potential in reducing computational complexity with complex multiplications replaced by bit-shifting arithmetic operations. However, most of existing methods resort to using full-precision hyper-parameters to avoid significant accuracy degradation, and so they achieve sub-optimal performance on the reduction of computation and memory cost. Therefore, we propose a novel efficient inference approach for resource-limited embedded devices. In our approach, integer-only PoT quantization scheme is designed to quantize both weights and activations into only PoT representations with fewer bit-width to achieve bit-wise inference. Additionally, a distribution-loss regularizer is developed to minimize the retraining disturbation introduced by quantizer. Furthermore, output weights are efficiently encoded in binary format by a two-stage compression pipeline. Finally, comprehensive experiments are conducted with different models and datasets. Experimental results indicate that our approach outperforms state-of-the-art approaches in terms of accuracy and compression rate, with more than 2 × the compression ratio previous methods achieve and 1.2 × accuracy improvement, under 3-bit precision with ResNet56.