Field Trial of Multi-Datacenter Distributed Training for LLM Based on Bandwidth Convergence and Two Parallel Strategies over 120km High-reliability 800Gbit/s C+L OTN

Yuyang Liu, Anxu Zhang, Xishuo Wang, Lipeng Feng, Kai Lv, Hao Liu, Xia Sheng, Xiaoli Huo, Junjie Li · 2025

We have conducted 175 billion parameters, 1024 GPUs large language model training with up to 99.41% (Pipeline parallel, PP) and 98.95% (Data parallel, DP) training efficiency in two distributed datacenters with an interconnection distance of 120km carried by 800Gbit/s C plus L WDM in the field-deployed high-reliability optical transport network.

Read the paper · More papers on PaperTik