Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation Fusion

H. Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou, Dazhao Cheng · 2025

The Mixture of Experts (MoE) architecture enhances model quality by scaling up model parameters. However, its development is hindered in distributed training scenarios due to significant communication overhead and expert load imbalance. Existing methods, which only allow for coarse-grained overlapping of communication and computation, slightly alleviate communication costs but at the same time, they introduce a notable impairment of computational efficiency. Furthermore, current approaches to addressing load imbalance often compromise model quality.

Read the paper · More papers on PaperTik