Text Classification by Bootstrapping with Keywords, EM and Shrinkage

Andrew McCallum, Kamal Nigam · 1999

When applying text classification to complex tasks, it is tedious and expensive to hand-label the large amounts of training data necessary for good performance. This paper presents an alternative approach to text classification that requires no labeled documents instead, it uses a small set of keywords per class, a class hierarchy and a large quantity of easily- obtained unlabeled documents. The key- words are used to assign approximate labels to the unlabeled documents by term- matching. These preliminary labels be- come the starting point for a bootstrap- ping process that learns a naive Bayes classifier using Expectation-Maximization and hierarchical shrinkage. When classifying a complex data set of computer science re- search papers into a 70-leaf topic hierarchy, the keywords alone provide 45% accuracy. The classifier learned by bootstrap- ping reaches 66% accuracy, a level close to human agreement.

Read the paper · More papers on PaperTik