Addressing the Heterogeneity of A Wide Area Network for DNNs
Hideaki Oguni, Kazuyuki Shudo · 2021
In general, deep neural networks (DNNs) achieve higher accuracy as the amount of training data increases. However, training data are often privacy sensitive, and they may not be collected. There are several methods that leave the training data decentralized in a wide area network and share models. These methods update the models locally based on stochastic gradient descent (SGD) and communicate to aggregate the models. The network bandwidth, training data, and machines are heterogeneous in a wide area network unlike distributed DNNs which use a computer cluster. Due to heterogeneity, the methods using synchronous communication, such as all-reduce SGD, are not suitable, and gossip SGD using asynchronous communication is a dominant method. In this paper, we show that when the network bandwidth is heterogeneous, conventional gossip SGD causes network congestion, and the learning efficiency is not greatly different from the case in which the network bandwidth is homogeneous. We show that the congestion problem can be solved by adjusting the communication frequency, that is, by training multiple times and communicating once. In many works, learning in local nodes and communicating with other nodes are alternated. Furthermore, we propose a warm-up technique to improve the learning efficiency. This proposed technique decreases the amount of communication with nodes that require a long communication time. We verify the effect of the proposed technique in experiments using CIFAR-10 and CIFAR-100.