Proactive learning: towards learning with multiple imperfect predictors
Jaime Carbonell, Pınar Dönmez · 2010
Label scarcity is a serious problem in many machine learning applications. In many domains such as classifying texts, images, etc., unlabeled data is readily available whereas labels are fairly expensive to obtain due to labeler availability, cost, and difficulty. Active learning is a paradigm that addresses this challenge by carefully selecting instances to he labeled. The goal is to improve the generalization performance of the learner with fewer labeling requests. Although active learning is well studied in the literature, it makes unrealistic assumptions. For instance, active learning assumes there is a unique omniscient oracle that works for free or charges uniform fee. In many real-world applications, it is quite possible there are multiple imperfect predictors with differing but unknown qualities. These qualities may vary from providing incorrect labels, failing to provide a label at all or charging non-uniform fees. The proactive learning paradigm addresses these problems to bridge the gap between active learning and more practical real-life scenarios. In this thesis, we first propose novel active learning methods for classification and rank learning problems that are shown to be quite effective in various real-world domains. We then describe a decision-theoretic framework that addresses learning with multiple predictors having non-ideal characteristics mentioned above with no apriori information, introducing proactive learning and how it can address various scenarios. Later, we focus on more specific aspects of proactive learning, especially coping with multiple fallible (noisy) predictors. We are interested in estimating the labeling accuracy of the predictors to select the most reliable ones in the absence of ground truth. We propose two novel approaches that achieve this goal when first the labeling accuracies are stationary and second when they vary with time. Our empirical evaluation demonstrates the ability to infer the predictor accuracy without any prior information in both synthetic and real-world datasets. Finally, we frame the problem as unsupervised risk estimation of multiple predictors as an alternative to the above approaches. The benefit is that it allows risk to he defined as a parametric function where the parameter governs the generation of the noisy predictor output. We propose maximum likelihood estimation framework as a solution. One of the most crucial consequences of this framework is to train classifiers that will minimize margin-based risk without using a single labeled example. The likelihood maximization framework yields statistically consistent estimators and hence effective estimation and training capabilities as supported by the thorough empirical evaluation.