On feature extraction by mutual information maximization

Kari Torkkola · IEEE International Conference on Acoustics Speech and Signal Processing · 2002

In order to learn discriminative feature transforms, we discuss mutual information between class labels and transformed features as a criterion. Instead of Shannon's definition we use measures based on Renyi entropy, which lends itself into an efficient implementation and an interpretation of “information potentials” and “information forces” induced by samples of data. This paper presents two routes towards practical usability of the method, especially aimed to large databases: The first is an on-line stochastic gradient algorithm, and the second is based on approximating class densities in the output space by Gaussian mixture models.

Read the paper · More papers on PaperTik