Exploitation of temporal structure in momentum-SGD for gradient compression
Tharindu B. Adikari, Stark C. Draper · 2021
Distributed optimization has become the norm for training machine learning models on large datasets. Learning bigger models on such systems leads to the exchange of large-volume updates. For this reason, limits on network communication in a distributed system can bottleneck learning progress. While compression techniques have been introduced to reduce bit-rates, no methods have yet leveraged the temporal structure that exist in consecutive vector updates. An important example is distributed momentum-SGD where temporal correlation is enhanced by the low-pass-filtering effect of applying momentum. In this paper we design methods that leverage temporal correlation to reduce bit-rates in systems employing momentum-SGD. We demonstrate that a significant reduction in the volume of communication can be realized. Experiments with the ImageNet dataset show that our proposed methods offers up to 40% bit savings compared to widely used methods such as Scaled-sign and Top-K.