Improving Generalization Performance of Adaptive Learning Rate by Switching from Block Diagonal Matrix Preconditioning to SGD
Yasutoshi Ida, Yasuhiro Fujiwara · 2020
Deep Neural Networks (DNNs) are widely used for various applications. Although adaptive learning rate algorithms are attractive for DNN training, their theoretical performance remains unclear. In fact, published analyses consider only simple optimization settings such as convex optimization, none of which are applicable to DNN training. This paper proposes TSO-ALRA, a two-stage optimizer using an adaptive learning rate algorithm; it is based on a full analysis of two approaches that do suit DNNs: parameter updates along geodesics on the statistical manifold and covariance structure of gradients. Our analysis reveals that the diagonal approximation used by existing adaptive learning rate algorithms inevitably degrades their efficiency. In addition, our analysis suggests that adaptive learning rate algorithms suffer drops in generalization performance in the last phase of training. To overcome these problems, TSO-ALRA combines an effective approximation technique and a switching strategy. Our experiments on several models and datasets show that TSO-ALRA efficiently converges with high generalization performance.