Improving Text Classification through Web Corpora

Rafael Guzmán-Cabrera, Manuel Montes-y-Gómez, Paolo Rosso, Luis Villaseñor-Pineda · 2007

A major difficulty with supervised approaches for text classification is that they require a great number of training instances (manually labeled examples) in order to construct an accurate classifier. In this paper we propose a semi-supervised method for text classification that is specially suited to work with very few training examples. This method consists of two main processes. The first one considers the automatic extraction of unlabeled examples from the Web. The second one focuses on the learning procedure, in particular, on the integration of the unlabeled examples into the original training set. Preliminary results indicate that our proposal can significantly improve the classification accuracy in scenarios where there are less than ten training examples available per class.

Read the paper · More papers on PaperTik