Unsupervised Document Classification with Informed Topic Models
Timothy A. Miller, Dmitriy Dligach, Guergana Savova · 2016
Document classification is an important and common application in natural language processing.Scaling classification approaches to many targets faces a bottleneck in acquiring gold standard labels.In this work, we develop and evaluate a method for using informed topic models to noisily label documents, creating a noisy but usable set of labels for training discriminative classifiers.We investigate multiple ways to train this noisy classifier, and the best performing method uses Wikipedia-seeded topic models to approximately label training instances without any supervision.We evaluate these methods on the classification task as well as in an active learning setting, in which they are shown to improve learning rates over traditional active learning.