WGLora:Efficient fine-tuning method integrating weights and gradient low-rank adaptation

Qingyun Lin, Wenlin He, qian Zhang, Zihan Peng, Zhendong Wu, Lilan Peng · 2024

To the ever-increasing size of weights and optimizer states. While existing low-rank adaptation techniques, like LoRA, reduce memory by introducing trainable low-rank matrices, they often fall short in matching the performance of full-parameter fine-tuning. To overcome this limitation, we introduce a novel fine-tuning approach that combines weight decomposition and gradient low-rank projection. By decomposing pre-trained weights into magnitude and direction components, and projecting gradients into a low-rank space, our method substantially reduces memory usage while maintaining gradient statistics. This approach enables efficient fine-tuning with reduced training and storage costs, without compromising model performance. When fine-tuning RoBERTa on GLUE tasks, our method achieves up to 63% reduction in optimizer state memory and around 21% decrease in GPU memory, while improving inference time by 12%. Notably, on the RTE dataset, our approach surpasses full-parameter fine-tuning in accuracy.

Read the paper · More papers on PaperTik