Aristos: Pipelining One-sided Communication in Distributed Mixture of Experts
Osayamen J Aimuyo · ACM SIGMETRICS Performance Evaluation Review · 2025
We propose Aristos, a communication-optimal, distributed algorithm that uses asynchronous communication interleaved with computation to specifically tackle the communication overhead of Distributed Mixture-of-Experts (DMoE) transformer models. DMoE, as implemented today, entails frequent synchronous All-to-All communication operations that scale linearly per step with a model's number of layers. Thus, we first seek clarification on how All-to-All bottlenecks DMoE and whether the interconnector communication algorithm is at fault. We investigate more than 21k All-to-All CUDA kernels and empirically confirm that their runtimes exhibit a significant long-tail distribution in both multi-node and single-node settings, respectively. We argue that this phenomenon is a shortcoming of the global barrier necessitated by the synchronous implementation of All-to-All in state-of-the-art collective libraries. We use these empirical insights to motivate Aristos, which obviates the global barrier, instead favoring asynchronous communication that yields native support for pipelining. Aristos also exposes tunable hyperparameters that navigate the trade-off between faster performance and reduced token dropping.