Towards Straggler-Tolerant and Accuracy-Aware Distributed DNN Training in Clouds

Shingo Okuno, Masahiro Miwa, Naoto Fukumoto · 2021

This study investigated how straggler mitigation affects accuracy during distributed training. While distributed training is one promising way to shorten training time, we cannot obtain a maximum performance improvement if some workers become stragglers due to a slowdown. We can avoid the performance impact from stragglers by excluding them from training. However, we face another problem of decreasing accuracy during training because gradients computed by excluded stragglers are not reflected in the training. To achieve both straggler tolerance and accuracy awareness, we should implement distributed training so that excluded stragglers return to the training once they recover from slowdown. Although some implementations supporting auto-scaling have been created, they have not evaluated the training accuracy during scale-in to exclude stragglers from training and scale-out to return them to training. Therefore, we conducted experiments to see how excluding stragglers causes accuracy degradation. We also discussed how the exclusion time and training data assignment can have an impact on accuracy during training.

Read the paper · More papers on PaperTik