A MoE-based Safety Fine-tuning Method for Multimodal Large Language Models

Runjia Zhang, Qi Wen Xue, Nanxin Zhang, Yu Wang, Xiurui Xie, Dongyang Zhang · 2025

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in understanding and reasoning across different modalities. Currently, the performance of the model on downstream tasks is mainly improved by various fine-tuning methods, however, traditional fine-tuning methods often lead to degradation of model safety, even when the fine-tuning process focuses on benign downstream tasks. This paper introduces an expert mixture-based safety fine-tuning method for multimodal models that simultaneously maintains model safety and enhances downstream task performance through multiple specialized expert adapters, coordinated by a learnable dynamic routing mechanism. Our approach incorporates eight LoRA adapter blocks in the vision encoder's attention layers, where four adapters are dedicated to safety control and four to downstream tasks. Experimental results on MiniCPM-V2.6, an 8B-parameter multimodal model, demonstrate that our method reduces the harmfulness rate from 33.08% to 30.45% compared to traditional LoRA. Notably, while conventional safety-oriented fine-tuning methods often sacrifice task performance, our approach achieves superior performance on downstream tasks, improving helpfulness on the TextVQA dataset by 19.44% compared to the base model. These results validate the effectiveness of our expert mixture approach in balancing safety and utility in multimodal models. Besides, we have also conducted domestic hardware platform adaptation experiments, contributing to the application of large models on Chinese domestic hardware platforms.

Read the paper · More papers on PaperTik