Quantization of Deep Neural Network Models Considering Per-Layer Computation Complexity for Efficient Execution in Multi-Precision Accelerators

Shen‐Fu Hsiao, Yu-Che Yen · 2021

Quantization of parameters and activation data in deep neural network (DNN) models plays an important role in multi-precision DNN hardware accelerators where the per-layer bit-widths are dynamically adjusted to speed up the computation. In this paper, we present an efficient quantization algorithm which considers the supported types of bit-width in the multi-precision DNN hardware when quantizing activations and parameters from the pre-trained models. Furthermore, we consider various computation complexity in different DNN layers in order to minimize the execution time in the DNN hardware. Experimental results show that the proposed algorithm has better quantization results for more efficient computation in multi-precision DNN hardware accelerators.

Read the paper · More papers on PaperTik