Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
Zeliang Zhang, Xiaodong Liu, Hao Cheng, Chenliang Xu, Jianfeng Gao · 2025
In this work, we address the memory overhead of deploying Mixture-of-Experts (MoE) architectures in Large Language Models (LLMs).While MoE layers improve LLM performance without increasing inference costs, the evergrowing number of experts inflates memory requirements, hindering practical deployment.Our empirical study reveals that some experts encode redundant knowledge during pretraining.We thus propose a method of grouping and pruning similar experts to improve the model's parameter efficiency.We validate the effectiveness of our method by pruning three state-of-the-art MoE architectures, including Mixtral, Deepseek-MoE, and Qwen.The evaluation shows that our method outperforms other model pruning methods on a range of natural language tasks.