CorrActive Learning: Learning from Noisy Data through Human Interaction

Ramesh M. Nallapati, Mihai Surdeanu, Christopher D. Manning · 2009

We introduce a new framework of supervised machine learning called CorrActive Learning, short for Corrective Active Learning. Similar to active learning, this setting involves learning through human interaction. However, unlike active learning which aims to acquire labels for unlabeled examples, corrActive learning addresses the problem where the set of training data provided to the supervised learner is noisy with respect to its labels. In this scenario, the objective is to accomplish the following two related goals simultaneously, using minimal assistance from the user: (a) clean up the noisy labeled data and (b) improve the performance of the supervised learner. As a solution, we present a simple algorithm that learns initially from the noisy labeled data, and proceeds to correct the labeling errors in the data iteratively by presenting to the user only those examples that are most likely to be mislabeled, and simultaneously learning from the corrected examples. Our preliminary experiments involving a human suggest that the new corrActive learner significantly improves the performance of a supervised classifier learned on noisy data. In addition, our synthetic experiments show that the corrActive learner is able to learn much faster than a learner that chooses examples using random sampling, when the labeling error rate is low to moderate (< 25%). At all error rates, the corrActive learner is able to identify the mislabeled examples much better than the random sampler. 1

Read the paper · More papers on PaperTik