Enhancing Embedded System Performance with an Adaptive Quantization Scheme for Binary Neural Networks
Wei Lu, Chaojie Yang, Shugang Zhang, Zhong Ma, Qin Yao, Wei Zheng · 2024
In order to deploy neural network models on resource-constrained embedded devices, floating-point neural network models need to be quantized into integer networks. Binary neural network is to quantize the activation/weights of the neural network model into 1 bit ({0,1} or {-1,1}). Since binary neural network saves a lot of storage cost and computational resource consumption, it is a promising technique for deploying deep neural networks on embedded devices. However, since there are only two possibilities for each parameter, binary neural network inevitably suffers from a severe loss of information, and optimising the binary neural network becomes difficult due to its discontinuities. The existing binary quantization studies have only been conducted at the level of simulated computational quantization, lacking validation on real hardware. In fact, the binary quantization methods that can really achieve significant acceleration effects on embedded devices by combining the device characteristics are of higher application value and research significance. In this paper, an adaptive quantization method for deploying binary neural networks on embedded systems is proposed, which can significantly improve the deployment accuracy of the binary neural network model on embedded systems.