Research on hyper parameter tuning for distributed optimization when data is sparse
Fang Xiao, Bao Maomao, Chen Yu-ting, Ting Mao · 2018
The inherent sequential updating mechanism of gradient descent is not conducive to asynchronous parallelism, because this way would induce much staleness gradient which could slow down model convergence rate. Thus, BSP (bulk synchronous parallel) model is generally adopted for distributed training. However, it would induce extra synchronization overhead which will greatly increase training time. Some researchers pointed out that when the data is sparse, the probability of staleness gradient occurred can be reduced obviously, and it is feasible to use fully asynchronous model in engineering. Based on previous theoretical research, this paper shows regret bound for distributed asynchronous training of FTRL-Proximal algorithm with adaptive learning rate, and analyzes the effect of learning-rate hyper-parameters and mini-batch size on model convergence rate.