Model Merging with Lightweight Fine-Tuning

Yan Pan, Yong Wang · 2024

Recently, model merging techniques have emerged as a solution for combining multiple models into a multi-task framework. However, merging often results in a significant drop in accuracy, necessitating fine-tuning to enhance performance. In this work, we refer to the performance of a model fine-tuned to convergence as the model potential, decided by which error basin model lies on when initialized. We identify that prior research on model merging has primarily focused on the performance of the merged model rather than its fine-tuned performance, and this oversight suggests that a better-performing merged model does not guarantee has better model potential. To address this issue, we propose a novel model merging method that facilitates better integration of knowledge from Model A and Model B, thereby enabling superior model potential. Specifically, we conceptualize the parameters of the merged model as a linear combination of those from Model A and Model B, allowing us to optimize the linear combination parameter (LCP) through training. To mitigate random fluctuations, we typically initialize the LCP at 0.5, ensuring that the initial merged model does not exhibit a preference for either Model A or Model B. Comprehensive experimental comparisons demonstrate that the features of Model A and Model B are more effectively combined with lightweight fine-tuning of the LCP.

Read the paper · More papers on PaperTik