Neural Networks for Unsupervised Learning Based on Information Theory
Jim Kay · 2000
Abstract We consider the statistical and information-theoretic basis of a class of artificial neural networks for unsupervised learning and processing. Taking a single processing unit as the basic component of the networks, we define information-theoretic objective functions and a class of activation functions, and derive the gradient-ascent learning rules. Several versions of the processor are considered, namely, the case in which the processor has a single binary output, the case in which the outputs are multinomial-winner-take-all, and that in which they are multivariate binary. In this latter case, local versions of the objective function are developed and these lead to local learning rules. Finally, the issue of computational complexity is briefly discussed, an alternative approach to the modelling is developed, and approximations are mentioned.