Most informative dimension reduction

Amir Globerson, Naftali Tishby · 2002

Finding effective low dimensional features from empir-ical co-occurrence data is one of the most fundamental problems in machine learning and complex data analy-sis. One principled approach to this problem is to rep-resent the data in low dimension with minimal loss of the information contained in the original data. In this paper we present a novel information theoretic princi-ple and algorithm for extracting low dimensional rep-resentations, or feature-vectors, that capture as much as possible of the mutual information between the vari-ables. Unlike previous work in this direction, here we do not cluster or quantize the variables, but rather ex-tract continuous feature functions directly from the co-occurrence matrix, using a converging iterative projec-tion algorithm. The obtained features serve, in a well defined way, as approximate sufficient statistics that capture the information in a joint sample of the vari-ables. Our approach is both simpler and more general than clustering or mixture models and is applicable to a wide range of problems, from document categorization to bioinformatics and analysis of neural codes.

Read the paper · More papers on PaperTik