Neuromanifolds and the Uniqueness of Relative Entropy

Stephen N. Winters-Hilt · 2021

This chapter discusses the theoretical detail to show modern arguments for the choice of relative entropy as difference measure on distributions. It shows a fundamentally derived variant of Expectation Maximization, referred to as “em”. The application of differential geometry methods to the study of statistical models traces back 1945, when it was noted that families of probability distributions could be described by a manifold and that the Fisher information matrix might be taken as a metric on that manifold. The chapter argues that the “simplest” divergence, the Kullback–Leibler divergence, is selected when maximizing log likelihood during learning, and that this is, fundamentally, because of the shortest path, or projection theorem. The information geometry methods described for families of probability distributions can just as easily be applied to neural networks; where the parameters are the connection weights.

Read the paper · More papers on PaperTik