Simpler Algorithm in Gradient Descent
David P. Woodruff, Youyu Zhu, Xiye Zhang · Frontiers in Educational Research · 2020
In Batch Gradient Descent, the most efficient constant learning rate should be , and L is the Lipschitz constant. In “Learning the Learning Rate”(Xiaoxia Wu et al., 2018), some of their algorithms seem t-o be inefficient. This paper is gonna to further improve their method. The searching-L algorithm is raised in the following passage, which limits the WNGrad-Batch’s complexity from square to linear by gradually approximating our learning rate to . In stochastic gradient descent, we find that b-y using the learning rate , which conforms the updating rule: , can have a more stable upper bound of the time complexity, which won’t have a parameter γ.