VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye, Li Lyna Zhang, Ting Cao, Cheng Li, Mao Yang · 2024

Scaling model size significantly challenges the deployment and inference of Large Language Models (LLMs).Due to the redundancy in LLM weights, recent research has focused on pushing weight-only quantization to extremely low-bit (even down to 2 bits).It reduces memory requirements, optimizes storage costs, and * Contribution during internship at Microsoft Research

Read the paper · More papers on PaperTik