Searching for High Accuracy Mix-Bitwidth Network Architecture

Haochong Lan, Zhiping Lin, Xiaoping Lai · 2020

The quantization of deep neural networks has drawn great research interest as well as industrial support for its tremendous advantages on boosting the inference speed of convolutional neural network models, especially on resource-constrained devices. Emerging research work and software framework for low-bit quantization, i.e., quantization under 8 bits, also shows strong potential. However, extreme low-bit neural network, e.g., binary neural network (BNN), suffers from significant accuracy loss. A previous work tried to improve the performance of extreme low-bit neural networks by adjusting channel numbers but achieves a poor balance between accuracy and model size. Alternatively, adjusting bitwidth for each layer gains a higher return on accuracy with an acceptable increase in model size. In this work, we employ a genetic algorithm to find the optimal bitwidth assignation for layers in VGG-Small and ResNet18. Both searched mix-bitwidth network architectures achieve higher accuracy with lower time and memory consumption, compared to their fix-bitwidth counterparts. A further investigation on the combination of BNN and the intermediate layer of higher bitwidth is also conducted.

Read the paper · More papers on PaperTik