Text classification with relatively small positive documents and unlabeled data

Fumiyo Fukumoto, Takeshi Yamamoto, Suguru Matsuyoshi, Yoshimi Suzuki · 2012

This paper addresses the problem of dealing with a collection of negative training documents which is suitable for relatively small number of positive documents, and presents a method for eliminating the need for manually collecting negative training documents based on supervised machine learning techniques. We applied an error correction technique to the results of negative training data obtained by the Positive Example Based Learning (PEBL). Moreover, we used a boosting technique to learn a set of negative data to train classifiers. The results using Japanese newspaper documents showed that the method contributes for reducing the cost of manual collection of negative training documents.

Read the paper · More papers on PaperTik