Data-Parallel Momentum Diagonal Empirical Fisher (DP-MDEF):Adaptive Gradient Method is Affected by Hessian Approximation and Multi-Class Data
Chenyuan Xu, Kosuke Haruki, Taiji Suzuki, Masahiro Ozawa, Kazuki Uematsu, Ryuji Sakai · 2022
Although data-parallel learning can efficiently reduce training time, the accuracy becomes poor when the batch size is extremely large. Adam, a prominent adaptive gradient optimizer which uses second-order information, is known to have better convergence than SGD, but its performance in large batch training highly depends on actual problems and the reasons for this remain unclear. While attempting answer to this, we found there are two factors playing import roles: the estimation accuracy of Hessian and the number of classes in data. Aiming to improve the former one, we introduce a new method called data-parallel momentum diagonal empirical Fisher (DP-MDEF), which collects the variance of the gradients between parallel workers for gradient adaptation and obtains superior precision compared to Adam and Momentum-SGD in large batch training. To discuss the latter one, we found that when the number of classes in a data-set is large, the variances of gradients form a long-tailed distribution, in which case gradient adaptive methods give bad solutions.