An Efficient Task-based All-Reduce for Machine Learning Applications

Zhenyu Li, James Davis, Stephen A. Jarvis · 2017

All-Reduce is a collective-combine operation frequently utilised in synchronous parameter updates in parallel machine learning algorithms. The performance of this operation - and subsequently of the algorithm itself - is heavily dependent on its implementation, configuration and on the supporting hardware on which it is run. Given the pivotal role of all-reduce, a failure in any of these regards will significantly impact the resulting scientific output.

Read the paper · More papers on PaperTik