Tailor : Datacenter Load Balancing For Accelerating Distributed Deep Learning
Zhengzhi Xu, Yifei Lu, Jingqi Li, Xu Ma, Lei Qian · 2022
Accelerating Distributed Deep Learning draws increasing attention recently. Existing solutions like gradient compression, pipelining computation/communication, flow scheduling based on gradient disparities, neglect load balancing for Distributed Deep Learning (DDL) flows in DataCenter Network. In this paper, we review the PS traffic pattern in usual DDL training and propose that minimizing tail Flow Completion Time (FCT) of DDL flows is objective for DDL traffic Load Balancing. Besides, traffics of DDL differ from traditional DataCenter in temporal and spatial distribution. Current DCN Load Balancing become unadaptable to optimization requirements of DDL traffic, so we propose Tailor, a brand new load balancing scheme for DDL traffics training in DCN. Tailor is a solution without complex switch modification that sufficiently leverages the known parameter amount of DDL flows to discover congestion and reroutes by ECMP-like hash function to compensate bandwidth for slower flows. With large-scale simulation, we validate the effectiveness of Tailor. Tailor behaves well in both symmetric and asymmetric topology. It reduces time cost of DDL training by decreasing about 20-30% tail FCT than other load balancing mechanisms.