Learning Word-to-Concept Mappings for Automatic Text Classification

Georgiana Ifrim, Martin Theobald, Gerhard Weikum, Luc De Raedt, Stefan Wrobel · 2005

For both classification and retrieval of natural language text documents, the standard document representation is a term vector where a term is simply a morphological normal form of the corresponding word. A potentially better approach would be to map every word onto a concept, the proper word sense and use this additional information in the learning process. In this paper we address the problem of automatically classifying natural language text documents. We investigate the effect of word to concept mappings and word sense disambiguation techniques on improving classification accuracy. We use the WordNet thesaurus as a background knowledge base and propose a generative language model approach to document classification. We show experimental results comparing the performance of our model with Naive Bayes and SVM classifiers.

Read the paper · More papers on PaperTik