Solving Distribution Shift in Quantization-Aware Training via Channel-Wise Standardization Distillation
Tianyang Cai, Junbiao Pang, Jiaxin Deng · 2024
With the continuous deepening of network models and the gradual increase of model parameters, model compression has become a necessary technology in the process of model deployment. Quantization of deep models can effectively reduce the computation and memory consumption based on low-bit parameters and bit-wise operations, and has recently attained much attention. However, due to quantization operations, quantized models are often affected by performance degradation. In this work, through extensive empirical analysis, we find that the performance degradation comes from the distribution shift between the Full Precision (FP) model and the quantized model. The distribution shift will lead to the continuous accumulation of quantization errors, which will affect the final performance of the model. In this paper, we address this issue from a distillation perspective to effectively eliminate such a shift. We present a simple and powerful method to eliminate the distribution shift via Channel-wise Standardization Distillation (CSD). CSD assumes that the quantized model and the FP model should have a similar distribution in the feature space. Our method generalizes well to different network architectures and various Quantization-Aware Training (QAT) methods. Extensive experiments demonstrate that our approach significantly outperforms the current State-Of-The-Art (SOTA) QAT methods and even FP counterparts.