UniFLE: Uniform Fusion of Multiple LoRA Experts for Backdoor Defense in Large Language Models
Shuai Zhao, Qika Lin, Yanhao Jia, X. H. Wu, Yuwen Li, Luu Anh Tuan · IEEE Transactions on Dependable and Secure Computing · 2026
Large language models (LLMs), which serve as a bridge between pre-training and task-specific adaptation, achieve state-of-the-art performance across several downstream tasks through full-parameter fine-tuning (FPFT). However, with the continuous growth of model parameter scales, FPFT requires substantial computational resources, which limits its practicality. Consequently, there is a growing shift toward parameter-efficient fine-tuning (PEFT) methods that update only a limited subset of model parameters, markedly reducing resource consumption. Although this paradigm fosters accelerated research advancements, it also introduces security risks. Empirical studies demonstrate that if LLM weights are backdoored, the backdoors can still be activated even after fine-tuning leveraging PEFT algorithms. To address the aforementioned issue, in this paper, we introduce a novel Uniform Fusion of multiple LoRA Experts algorithm, named UniFLE, designed to defend against backdoor attacks. Specifically, the UniFLE algorithm pioneers the insertion of multiple LoRA experts into the MLP blocks to expand the updatable feature subspace during fine-tuning, and fuses all experts to decouple backdoor features. Additionally, to enhance the diversity of LoRA experts, we introduce a diversity regularization loss that constrains correlations between different experts and encourages them to learn more novel features. The theoretical analysis demonstrates that the UniFLE algorithm effectively decreases the mutual information between the model's intermediate representations and the backdoor representations. To validate the effectiveness of the UniFLE algorithm, we conduct experiments across four tasks, four state-of-the-art LLMs, and three backdoor attack methods. The results consistently demonstrate that our UniFLE algorithm effectively defends against backdoor attacks while preserving model performance. We aspire for our method to enhance model security and contribute to the advancement of the LLM community.