Semi-automatic classification: using active learning for efficient class coverage
Filipe Rhodes da Fonseca, Nuno Escudeiro · 2012
In some classification tasks, such as those related to the automatic building and maintenance of text resources, it is expensive to obtain labeled instances to train a classifier although it is common to have massive amounts of data available at low cost. Unlike supervised learning, that requires a fully pre-labeled training set, active learning allows asking an oracle to label only the most informative instances given the specific purpose of the learning task and the available data. Moreover, active learning is an iterative process that may be halted when the potential utility of the unlabeled instances remaining in the working set is low. Active learning generally requires a lower labeling effort to build accurate classifiers than supervised learning. However, common active learning approaches assume the availability of a pre-labeled set, covering all the target classes, to initialize the learning process. The labeling effort required to build this initialization set is not generally considered when analyzing the performance of the learning process. When in presence of imbalanced class distributions, identifying labeled instances from minority classes might be very demanding, requiring extensive labeling, if queries are randomly selected. Nevertheless, these minority classes are the most critical to certain classification tasks, such as, detection of fiscal fraud and rare diseases diagnosis. In such circumstances, evaluating the performance and building a classifier based exclusively in accuracy might not be appropriate since an accurate classifier might still fail to identify minority classes – the critical ones – with a little impact in accuracy. A novel approach to active learning is required in order to comply with these cases. Besides accuracy, it is also important to assure that the classifier being built is aware of all target classes irrespectively of their distribution. It is our belief that it is possible to develop an active learning strategy that builds accurate classifiers being aware of all the target classes at a reduced labeling effort – that is, at low cost – when compared to current approaches. In this thesis we propose a strategy for active learning that comprises an active learning criterion to select queries and a stopping criterion to halt the learning process when the utility of the remaining unlabeled instances is low. D-Confidence, our query selection approach, is based on a query selection criterion that aggregates the posterior classifier confidence and the distance between unlabeled instances and known classes. This criterion is biased towards instances belonging to unknown classes – low confidence – that are located in unexplored regions in the input space – high distance to known classes. The stopping criterion in our strategy, hcw, is an ensemble of classification gradient and steady entropy mean, two base indicators of the utility of unlabeled instances. Classification gradient provides evidence on the differences of the predicted labels between two consecutive iterations of the learning process. Steady entropy mean provides information on the stability of the distribution of the entropy of predictions between two consecutive iterations. This strategy is expected to identify exemplary instances from all the target classes, independently of their frequency, being able to train an accurate classifier while requiring a reduced labeling effort when compared to common active learning approaches.