Learning the Similarity of Documents: An Information-Geometric Approach to Document Retrieval and Categorization

Thomas Frank Hofmann · 1999

The project pursued in this paper is to develop from rst information-geometric principles a general method for learning the similarity between text documents. Each individual document is modeled as a memoryless information source. Based on a latent class decomposition of the term-document matrix, a lowdimensional (curved) multinomial subfamily is learned. From this model a canonical similarity function { known as the Fisher kernel { is derived. Our approach can be applied for unsupervised and supervised learning problems alike. This in particular covers interesting cases where both, labeled and unlabeled data are available. Experiments in automated indexing and text categorization verify the advantages of the proposed method. 1 Introduction The computer-based analysis and organization of large document repositories is one of today's great challenges in machine learning, a key problem being the quantitative assessment of document similarities. A reliable similarity measure...

Read the paper · More papers on PaperTik