Structured Optimal Brain Pruning for Large Language Models
Jiateng Wei, Quan Lü, Ning Jiang, Siqi Li, Jingyang Xiang, Jun Chen, Yong Liu · 2024
The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs).Network pruning provides a practical solution to this problem.However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning.The former relies on special hardware to accelerate computation, while the latter may need substantial computational resources.In this paper, we introduce a retraining-free structured pruning method called SoBP (Structured Optimal Brain Pruning).It leverages global first-order information to select pruning structures, then refines them with a local greedy approach, and finally adopts module-wise reconstruction to mitigate information loss.We assess the effectiveness of SoBP across 14 models from 3 LLM families on 8 distinct datasets.Experimental results demonstrate that SoBP outperforms current state-of-the-art methods.