Mast: Efficient Training of Mixture-of-Experts Transformers with Task Pipelining and Ordering
Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang, Bo Li, Xiaowen Chu · 2025
The utilization of the sparsely activated mixture-of-experts (MoE) technique has enabled the expansion of modern large language models (LLMs) to trillion-level sizes while maintaining a sub-linear increase in computations. This involves equipping an MoE layer with multiple experts, where only one or two experts are activated for each input data. However, the dynamic activation of MoE experts introduces extensive communications, limiting the scaling efficiency of distributed systems. In this work, we propose Mast to efficiently train MoE models by pipelining and re-ordering communication and computation tasks to effectively hide communication costs. Specifically, we first propose to overlap tasks in both attention layers and MoE layers. Then we theoretically analyze the task overlaps between communications and computations, identifying the inefficiencies of existing schedules. We then develop an optimization formulation to determine a near-optimal order for task pipelining with the objective of minimizing iteration time. We conduct extensive experiments on two 32-GPU clusters employing 432 configured MoE layers and three real-world MoE models based on BERT, GPT-2 and Mistral. The experimental results demonstrate that Mast outperforms state-of-the-art MoE training systems (DeepSpeed-MoE, Tutel, PipeMoE and CoCoNet) with an average speedup 1.13 ×-1.43 × on the MoE models.