The Learning Rate for The Gradient Descent

Settawut Tanakam, Thap Panitanarak · 2022

To minimize the continuous and differentiable objective function f, the gradient descent, a first-order iterative method, can provide a sequence number {x_i} that converges to the minimum. The iteration can be written as x_(i+1) = x_i - t ∇f_i where t is a learning rate and i is an iterative number. However, the value of the learning rate affects the whole algorithm. If the value is too small, the sequence {x_i} will slowly converge, and if the value is too large, the sequence {x_i} will oscillate or diverge. To prevent this problem, the learning rate t is adjusted for each iteration i. Therefore, the problem is reduced to minimize f ̃(t_i ) = f(x_i - t_i ∇f_i). The line search method will choose the optimal t_i be the learning rate in iteration i. In this project, we improve the backtracking line search by choosing a smaller learning rate for gradient descent based on the previous iteration. After the experiment, the modified method performs as well as the standard method for simple functions and better in more complex functions in terms of running time and number of iterations.

Read the paper · More papers on PaperTik