Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts Models

Juncai Liu, Jessie Hui Wang, Yimin Jiang · 2023

Scaling models to large sizes to improve performance has led a trend in deep learning, and sparsely activated Mixture-of-Expert (MoE) is a promising architecture to scale models. However, training MoE models in existing systems is expensive, mainly due to the All-to-All communication between layers.

Read the paper · More papers on PaperTik