Gradient Tracking with Multiple Local SGD for Decentralized Non-Convex Learning
Songyang Ge, Tsung‐Hui Chang · 2023
The stochastic Gradient Tracking (GT) method for distributed optimization, is known to be robust against the inter-client variance caused by data heterogeneity. However, the stochastic GT method can be communication-intensive, requiring a large number of communication rounds of message exchange for convergence. To address this challenge, this paper proposes a new communication-efficient stochastic GT algorithm called the Local Stochastic GT(LSGT) algorithm, which adopts the local stochastic gradient descent (local SGD) technique in the GT method. With LSGT, each agent can perform multiple SGD updates locally within each communication round. Although it is not known previously whether the stochastic GT method can benefit from the local SGD, we establish the conditions under which our proposed LSGT algorithm enjoys the linear speedup brought by local SGD. Compared with the existing work, our analysis requires less restrictive conditions on the mixing matrix and algorithm stepsize. Moreover, it reveals that the local SGD does not only reserve the resilience of the stochastic GT method against the data heterogeneity but also speeds up reducing the tracking error reduction in the optimization process. The experimental results demonstrate that the proposed LSGT exhibits improved convergence speed and robust performance in various heterogeneous environments.