AutoMPQ: Automatic Mixed-Precision Neural Network Search via Few-Shot Quantization Adapter
Ke Xu, Xiangyang Shao, Ye Tian, Shangshang Yang, Xingyi Zhang · IEEE Transactions on Emerging Topics in Computational Intelligence · 2024
Model quantization has gained significant attention as a widely adopted technique for compressing deep neural networks. Recent advancements in network quantization have achieved remarkable outcomes by employing mixed-precision quantization, which dynamically adjusts the bit-width for each layer based on its specific requirements. Although genetic algorithms offer a viable approach for implementing mixed-precision search, the high cost of evaluation poses a significant challenge to the efficiency of the search process. This challenge arises from the extensive time required for fine-tuning and the inter-dependencies among different layers. To alleviate this bottleneck, we propose AutoMPQ, an automatic mixed-precision neural network search framework. AutoMPQ introduces an innovative evaluation mechanism based on a few-shot quantization adapter strategy. This approach significantly reduces the evaluation cost by efficiently tuning the meta-parameters of batch normalization (BN), mixed-precision convolution (MPConv), and mixed-precision ReLU (MPReLU) using minimal calibration data. In the fine-tuning stage, we validate the effectiveness of the knowledge distillation strategy guided by the full-precision network, leading to competitive results. Extensive experiments conducted on the CIFAR-10 and ImageNet datasets demonstrate the superiority of AutoMPQ over handcrafted uniform bit-width counterparts and other existing mixed-precision techniques.