Rapid Accuracy Loss Evaluation for Quantized Neural Networks Using Attention-Based Multifactor Modeling
Wei Lu, Chaojie Yang, Zhong Ma, Qin Yao · 2024
Neural network quantization, the process of converting floating-point models into low-bit-width integer models, has become a key technology for reducing the computational and storage costs of neural network models. However, quantization inevitably leads to a loss of accuracy, making it particularly important to assess this loss quickly and accurately. Rapid evaluation can significantly save time and resources, especially during the model development and deployment phases, and facilitates rapid model iteration and optimisation. However, quantization accuracy loss is influenced by multiple factors, often involving multimodal data, which are difficult to model and integrate effectively. This paper proposes a multifactor modeling method based on attention mechanisms, which aims to analyse and simulate the relationship between quantization factors and neural network performance, thereby providing a method to quickly and accurately assess post-training quantization accuracy loss. Experimental results show that this method can quickly predict the accuracy of different neural network models on various computational tasks without relying on extensive test data, with a prediction error of only about 2%, effectively supporting rapid model development and iterative optimization. Codes will be available on https://github.com/ycjcy/Accuracy-Loss-Evaluation.