Gradients on matrix manifolds and their chain rule

Fabian Joachim Theis · University of Regensburg Publication Server (University of Regensburg) · 2005

Optimization on matrix manifolds is an important tool in machine learning and neural networks, and for local algorithms it is often necessary to compute the gradient of a function defined on a matrix space. Recent advances in information geometry have shown that by adopting to the inherent geometry of the search space, convergence speed can be greatly increased. Here, we give an overview of various gradient calculations on the group Gl(n) of all invertible (n £ n)-matrices. We review how to introduce a right-invariant, so-called natural metric on Gl(n). Then we calculate the resulting, natural gradient of a function on Gl(n) in terms of the ordinary, Euclidean gradient and present a chain rule for gradients, which is illustrated by two examples. We finish with generalizations to over- and undercomplete cases, realized by a semidirect product.

Read the paper · More papers on PaperTik