Dynamic Low-Rank Adaptation Based Pruning Algorithm for Large Language Models
Linfeng Li, Lei Guo · 2024
Deep neural networks have been getting better and better performance while the parameter size has become larger and larger, and have now entered the era of large language models (LLM). However, large language models with huge number of parameters are difficult to be deployed on edge devices with limited computing power and memory resources. Network pruning techniques can reduce the storage and computation requirements of large language models. Although Dynamic Low-Rank Adaptation (DyLoRA) allows efficient fine-tuning of LLM, the deployment of LLM is still hindered by large model sizes and computational costs. In this paper, we propose a DyLoRA-based pruning algorithm that uses gradient updating of DyLoRA module to define the importance of parameters, and then prune the least important parameters based on the resulting importance scores. Through experiments on commonly used evaluation datasets, it is demonstrated that the proposed pruning algorithm is able to achieve 96.60% of the performance of the original model with 20% of the parameters pruned, which outperforms the previous pruning algorithms.