Exploring Quantization Techniques for Large-Scale Language Models: Methods, Challenges and Future Directions
Ao Shen, Zhiquan Lai, Dongsheng Li · 2024
Breakthroughs in natural language processing (NLP) by large-scale language models (LLMs) have led to superior performance in multilingual tasks such as translation, summarization, and Q&A. However, the size and complexity of these models raise challenges in terms of computational requirements, memory usage, and energy consumption. Quantization strategies, as a type of model compression technique, have gained attention for their advantages in reducing model size and accelerating inference speed. In this paper, we review the rapid development of quantization techniques for LLMs, and systematically explore methods such as post-training quantization (PTQ), quantization-aware fine-tuning (QAF), and quantization-aware training (QAT), which provide a comprehensive solution to the resource-intensive problem of LLMs from training to post-deployment. We also analyze state-of-the-art benchmarks and datasets to evaluate the effectiveness of quantitative methods in terms of performance retention and computational efficiency. The purpose of this review is to provide researchers with a snapshot of the latest progress in LLM quantization techniques, helping them to quickly grasp the dynamics of the field and understand the key techniques and challenges, so that they can more efficiently devote themselves to this evolving research area.