Understanding and Characterizing Communication Characteristics for Distributed Transformer Models

Quentin G. Anthony, Benjamin Michalowicz, Jacob Hatef, Lang Xu, Mustafa Abduljabbar, Aamir Shafi, Hari Subramoni, Dhabaleswar K. DK Panda · IEEE Micro · 2025

The transformer architecture has revolutionized many applications, such as large language models. This progress has been largely enabled by distributed training, yet communication remains a significant bottleneck. This article examines the communication behavior of transformer models, focusing on how different parallelism schemes in multinode/multi-GPU training communicate data. We use Generative Pre-trained Transformer-based language models as a case study due to their prevalence. We validate our empirical results using analytical models. Our analysis reveals practical insights and potential areas for further optimization in framework and high-performance computing middleware design.

Read the paper · More papers on PaperTik