Deep Learning Framework with Arbitrary Numerical Precision
Masato Kiyama, Motoki Amagasaki, Masahiro Iida · 2019
Deep neural networks (DNNs) have recently shown outstanding performance in solving problems in many domains. However, it is difficult to run such applications on mobile devices due to limited hardware resources. Quantization is one method to reduce the hardware requirements. By default, 32-bit floating-point numbers are used in DNNs, while quantization uses fewer bits, such as 4-bit fixed points, at the cost of precision. Previous research has explored two problems related to this: (1) differences between software emulation and implementation that affect model accuracy and (2) lowered accuracy during normalization. In this paper, we developed a new DNNs framework, PyParch, that allows easy manipulation of quantization and propose a training method for fitting to a hardware-friendly model. We show that our developed tool can solve the two problems mentioned above. Quantized models described in previous methods need 18 bits in order to recover the original accuracy, whereas our method requires only 14 bits.