T3P: Topology-Tailored Tensor Parallelism
Saar Ben Yochana, Chen Avin, Gabriel Scalosub · 2025
As deep learning models continue to grow in scale and complexity, methods of distributed machine learning training, and particularly those used for large language models (LLMs), have become a critical ingredient in making such computations efficient and feasible. In such contexts, tensor parallelism (TP) is widely employed to distribute computations across multiple accelerators. However, since TP mandates frequent and high-volume communication between devices, the underlying network characteristics significantly influence performance. Previous work was mostly either model-agnostic or topology-agnostic and did not pick provably optimal configurations. This study presents Topology-Tailored Tensor Parallelism, T3P, an efficient algorithm that identifies the communication-optimal TP sharding configuration (within the considered search space) based on both the model architecture and the network topology. In particular, we show that T3P is optimal for any given resharding cost model.