A Quantized Neural Network Library for Proper Implementation of Hardware Emulation
Masato Kiyama, Yasuhiro Nakahara, Motoki Amagasaki, Masahiro Iida · 2019
Deep neural networks (DNNs) have recently shown outstanding performance in many domains. However, running DNN applications on mobile devices with limited hardware resources is difficult because DNN requires heavy computations. Quantization is one of network compression techniques to reduce the hardware requirements because it uses fewer bits, such as the 8-bit fixed points instead of 32-bit floating-point numbers. DNN libraries use floating-point numbers after quantization because they assume that all data to be calculated in floating-point format. Therefore, a mismatch occurs in running hardware emulation. In this paper, we propose a method of calculation for hardware emulation and developed a new DNNs library based on our method. We show that our library is capable of exact hardware emulation.