Anytime Minibatch with Stale Gradients : (Invited Paper)
Haider Al-Lawati, Nuwan S. Ferdinand, Stark C. Draper · 2019
Large-scale machine learning problems are often solved using distributed optimization techniques by dividing tasks across multiple compute nodes. Due to variability in compute time amongst nodes, distributed systems suffer from slow nodes known as stragglers that can have a big impact on convergence speed. Recently, Anytime Minibatch (AMB) technique has been proposed to speed up convergence by exploiting work done by stragglers rather than avoiding them. In AMB, workers are given a fixed time to calculate gradients, followed by a fixed communication time to average gradients. We observe that nodes stay idle during communication time. In our work, we allow workers to compute gradients during communication time resulting in what is known as stale gradients due to gradients calculated on out-of-date parameters. Our simulation results show that such approach achieves up to 60% faster convergence in wall time.