Efficient on-line learning with diagonal approximation of loss function Hessian

Paweł Wawrzyński · 2019

The subject of this paper is stochastic optimization as a tool for on-line learning. New ingredients are introduced to Nesterov's Accelerated Gradient that increase efficiency of this algorithm and determine its parameters that are otherwise tuned manually: step-size and momentum decay factor. In this order a diagonal approximation of the Hessian of the loss function is estimated. In the experimental study the approach is applied to various types of neural networks, deep ones among others.

Read the paper · More papers on PaperTik