Scaling deep learning on GPU and knights landing clusters
Yang You, Aydın Buluç, James Weldon Demmel · 2017
Training neural networks has become a big bottleneck. For example, training ImageNet dataset on one Nvidia K20 GPU needs 21 days. To speed up the training process, the current deep learning systems heavily rely on the hardware accelerators. However, these accelerators have limited on-chip memory compared with CPUs.