FP4-Quantization: Lossless 4bit Quantization for Large Language Models

Jie Wang, Huanxi Liu, Dawei Feng, Jie Ding, Bo Ding · 2024

Large language models(LLMs) have demonstrated exceptional performance across a wide range of tasks. However, their extensive computational and storage requirements hinder their widespread deployment. To address this, low-bit quantization has emerged as a highly effective approach to reducing the inference cost of LLMs. Nevertheless, the existing repertoire of 4-bit quantization techniques is plagued by a substantial decline in model precision. In this paper, we introduce a novel 4-bit weight quantization method, FP4-Quantization, which leverages a 4-bit floating-point(FP4) representation that aligns better with the weight distribution characteristics of LLMs. Furthermore, it incorporates a Low-Rank Quantization Error Correction(LREC), involving progressive fine-tuning of low-rank parameters to rectify quantization errors, thereby enabling the achievement of precision-preserving 4-bit weight-only quantization. Our Experimental results on multiple zero-shot tasks demonstrate that FP4-Quantization achieves 4-bit weight quantization with an accuracy degradation of less than 0.5%.

Read the paper · More papers on PaperTik