Diffindo: Accelerating Distributed GANs with Auxiliary Generators and Discriminators
Xiaoming Han, Boan Liu, Dazhao Cheng · 2024
Integrating edge computing with Generative Adversarial Networks (GANs) leads to significant communication overhead in centralized federated learning systems due to frequent synchronization between generators and discriminators, as well as the transmission of large generated samples. In addition, this close coupling can cause excessive GPU memory consumption, especially with coarse-grained deployment strategies, resulting in memory thrashing and reduced training speeds. To tackle these issues, we present Diffindo, a novel distributed GAN training system that reduces training time and optimizes GPU memory allocation while maintaining accuracy. By utilizing fine-grained deployment strategies and developing innovative task scheduling algorithms for servers and workers, we enhance training efficiency. Additionally, we implement a computation-communication overlapping strategy to improve resource utilization. Experimental results show that Diffindo outperforms state-of-the-art GAN training systems, achieving 13% higher accuracy and 32% faster training speeds.