Quantization of Deep Neural Networks for Improving the Generalization Capability

신성호 · Seoul National University Open Repository (Seoul National University) · 2020

Deep neural networks (DNNs) achieve state-of-the-art performance for various applications such as image recognition and speech synthesis across different fields.However, their implementation in embedded systems is difficult owing to the large number of associated parameters and high computational costs.In general, DNNs operate well using low-precision parameters because they mimic the operation of human neurons; therefore, quantization of DNNs could further improve their operational performance.In many applications, word-length larger than 8 bits leads to DNN performance comparable to that of a full-precision model; however, shorter word-length such as those of 1 or 2 bits can result in significant performance degradation.To alleviate this problem, complex quantization methods implemented via asymmetric or adaptive quantizers have been employed in previous works.In contrast, in this study, we propose a different approach for quantization of DNNs.In particular, we focus on improving the generalization capability of quantized DNNs (QDNNs) instead of employing complex quantizers.To this end, first, we analyze the performance characteristics of quantized DNNs using a retraining algorithm; we employ layer-wise sensitivity analysis to investigate the quantization characteristics of each layer.In addition, we analyze the differences in QDNN performance for different quantized network sizes.Based on our analyses, two simple quantization training techniques, namely adaptive step size retraining and gradual quantization are proposed.Furthermore, a new training scheme for QDNNs is proposed, which is referred to as high-low-high-low-precision (HLHLp) training scheme, that allows the network to achieve flat minima on its loss surface with the aid of quantization noise.As the name suggests, the proposed training method employs high-low-high-low precision for network training in an alternating manner.Accordingly, the learning rate is also abruptly changed at each stage.Our obtained analysis results include that the i

Read the paper · More papers on PaperTik