Gossip Distillation: Decentralized Deep Learning Transmitting Neither Training Data Nor Models
Taisuke Moriwaki, Kazuyuki Shudo · 2023
While deep learning requires a large amount of training data to obtain a highly accurate model, it is not always possible to collect the data in one place for privacy and other reasons. Furthermore, eliminating centralized servers and allowing all nodes to communicate in a decentralized way, improves fault tolerance and eliminates the unfairness of servers getting learned models first. In existing methods, a node transmits a learned model to other nodes. Our proposal, Gossip Distillation is to apply Knowledge Distillation to communication between nodes. A node transmits inference results on common data, not the model itself. The method reduces the amount of communication and makes it possible to combine multiple different models for each node. For CINIC-10, only 5.18 MiB of the inference results are transmitted between nodes though existing methods requires 49.03 MiB of ResNet-18 as the main model to be transmitted. In addition, the proposed method allows different sub models to be trained in parallel with the main model. Achieved accuracies are comparable with existing centralized methods.