Continual Defense Against Evolving Jailbreaks: A Multi-Agent Adversarial Framework with Linear Gating MoE
Qianqiao Xu, Feng Liu · 2025
Large Language Models (LLMs) are vulnerable to deceptive jailbreak attacks inducing harmful outputs. Existing defenses suffer from catastrophic forgetting in continual defense learning against evolving attacks. To address the limited space issue caused by defenders attempting to reduce defense forgetting when learning long sequence tasks, we propose a Mixture-of-Experts (MoE) based continual defense learning method, which only activates the most suitable experts to learn new tasks and freezes the other experts to prevent forgetting the old defense tasks. Our approach employs in-context learning (ICL) to simulate an attack-defense LLM-based multi-agent framework, where an attacker generates sequential novel attacks, a defender continuously learns to counter the attacker, and an evaluator assesses the defense. Crucially, we introduce a linear gating MoE continual learning method for the defender. This method uses a linear gate to assign new attack tasks to specific expert networks (formed by partitioning model neurons), isolating training to prevent forgetting of established defenses. Experiments demonstrate that our framework significantly outperforms baselines in continual defense learning, achieving higher defense success rates across sequentially updated attack tasks.