Using the Natural Gradient for training Restricted Boltzmann Machines
Benjamin Schwehn · 2010
Restricted Boltzmann Machines are difficult to train: Only an approximation of the gradient of the objective function is available. The objective function itself is intractable. Therefore most of the efficient convex optimisation techniques such as conjugate gradient descent or line search methods cannot be used. Using second order information is also not easy: evaluating the Hessian or an approximation of the Hessian is hard, and even if the Hessian is available, the dimensionality is usually too high to allow for inversion of the Hessian. Therefore RBMs are usually trained by stochastic gradient ascent, using a rough approximation of the gradient obtained by Gibbs sampling. It is known that learning by stochastic gradient ascent can be improved by including the covariance between stochastic gradients. However, this requires inverting a covariance matrix of size proportional to the dimensionality of the gradient. Recently, a new efficient approximate natural gradient method has been proposed. This method has shown to significantly improve training of ANNs, but hasn’t been tried on RBMs yet. In this report I will show that natural learning can indeed improve the training of RBMs during the convergence phase. However, natural learning methods can lead to instabilities and very slow initial learning. If natural learning is to be used, it should not be used for the initial training phase. I will also describe