Optimizing Stochastic Gradient Descent Using the Angle Between Gradients

Chongya Song, Alexander Pons, Kang K. Yen · 2020

In the field of machine learning, Stochastic Gradient Descent has been proven an effective method to shorten the time spent on minimizing the output cost. Due to the fact that the data pattern of each mini-batch could be somewhat varied from the full dataset, most existing optimization algorithms attempt to alleviate this variance via computing certain calibration terms associated with the previous gradient. Since the previous gradient is computed from the past Stochastic Gradient Descent state, it can adversely affect the calibration terms introducing deviation that result in an inaccurate new gradient. To resolve this problem, we propose a method that reduces the aforementioned deviation via applying a preprocessing technique to the previous gradient prior to its usage. The technique uses the angle between the previous and the current gradients to improve the precision of the calibration terms, reducing the effects of the deviations. Empirical results are obtained from incorporating the proposed method with a fully-connected vanilla neural network. The proposed technique is evaluated using the MNIST dataset against 10 previously proposed optimization algorithms. The experiment shows the benefits in reducing the cost result from adopting the proposed preprocessing technique to improve the gradient derivation.

Read the paper · More papers on PaperTik