Post-Training Quantization for Longformer with Chunkwise Quantization Granularity and Optimized Percentile
Qibin Chen, Yijin Teng, Hui Zhang, Kai Jiang, Qiang Duan, Xue Li, Xinxin Zhao, Rui Li · 2022 7th International Conference on Computer and Communication Systems (ICCCS) · 2022
Transformer-based models have been verified successful in many natural language processing and computer vision tasks. Because of computational complexity, many efficient transformer variants have been proposed, including the Long-former, which aims for long document processing. In this paper, we present an effective post-training quantization scheme for Longformer. Based on sliding window attention in Longformer, we propose chunkwise quantization. It can decrease quantization noise caused by significant gaps between ranges of different windows. Besides, to reduce quantization noise caused by clipping, we optimize percentile value by minimizing mean squared error between the original and quantized matrixes. The quantization scheme is evaluated on the TriviaQA task, and the performance is comparable to the float32 model. In addition, it is important that the quantization scheme can be extended to other efficient transformer-based models.