Bag-of-Concepts Document Representation for Bayesian Text Classification

Marcos Mouriño-García, Roberto Perez-Rodriguez, Luis Anido, Miguel Gómez-Carballa · 2016

The classification of text documents into a number of pre-defined categories has many application scenarios, for example the classification of news items into thematic sections. Documents to be classified are commonly represented by a bag-of-words feature vector. The bag-of-words model cannot handle two language phenomena: synonymy and polysemy, besides, dimensions of feature vectors are orthogonal. In order to effectively address those problems, some researchers adopt a bag-of-concepts representation of documents-understanding concept as "unit of meaning'". This paper reports a comprehensive experimental evaluation of the efficiency of a bag-of-concepts representation for Bayesian text classification, tackling synonymy and polysemy, and exploiting semantic relatedness between concepts to alleviate the problem of orthogonality-following an approach that we call semantic expansion. Results of experiments performed on three corpora widely used as benchmarks-Reuters, OHSUMED, and 20Newsgroups-show that: the efficiency of the bag-of-concepts approach is very dependent on the capacity of the semantic annotator for extracting concepts and on the characteristics of particular corpora, peaking on OHSUMED, and that it performs especially well when the number of training samples is small. In particular, for the shorter training sequence bag-of-concepts outperforms bag-of-words by 43.67% in OHSUMED and 22.44% in 20Newsgroups. This work provides useful insights to researchers that aim at applying bag-of-concepts representations, for example for organizing scientific articles in accordance with their thematic.

Read the paper · More papers on PaperTik