DQMix-BERT: Distillation-aware Quantization with Mixed Precision for BERT Compression
Yan Tan, Lei Jiang, Peng Chen, Chaodong Tong · 2023
Transformer-based architecture models like BERT have performed excellently for various Natural Language Processing (NLP) tasks. However, these models are usually computationally expensive with a large number of parameters. As a result, deploying them in edge devices has become a challenging task. The existing compression work on lower-precision quantization still has a severe accuracy decrease and rarely focuses on the information hidden in the different modules of the model. In this paper, we propose a distillation-aware quantization with mixed precision method combined with quantization and knowledge distillation. We achieve the ultra-low mixed precision quantization with the different sensitivity of different modules of BERT. Moreover, we leverage knowledge distillation to reduce the model accuracy degradation. We extensively test our method on four GLUE tasks. It shows that DQMix-BERT outperforms the other BERT compression methods and even achieves comparable performance to the original BERT model while achieving ~8x compression.