SILoR: SVD-Based Initialization for Low-Rank Adaptation in Vision Transformers
Yue Zhu, Huan Liu, ShuXin Chen · 2025
Fine-tuning pre-trained vision models for downstream tasks is cornerstone of computer vision, enabling the adaptation of general-purpose models to specialized applications. However, the growing size and complexity of models like Vision Transformers(ViTs) make traditional fine-tuning increasingly impractical due to its high computational and storage costs. Parameter-efficient fine-tuning(PEFT) methods, such as LowRank Adaptation(LoRA), address this by introducing lightweight, trainable modules while freezing most pre-trained parameters. LoRA leverages low-rank decomposition to achieve efficient finetuning with minimal additional parameters, reducing resource consumption while maintaining competitive performance. Despite its advantage, LoRA’s initialization often limits its representation capabilities and causes misalignment with pre-trained features. To address this, we propose SILoR, a novel initialization method for LoRA that uses singular value decomposition(SVD) of pretrained weights. SILoR align LoRA’s representation subspace more compact with the original feature space of the pre-trained model, enabling flexible and balanced feature aggregation. Extensive experiments across diverse datasets shwo that SILoR outperforms other PETL methods and even surpasses full finetuning in many cases, all with negligible inference overhead due to its efficient re-parameterization design.