Reduction methods for multi-label classification
Cheng Li · 2019
Multi-label classification is an important machine learning task wherein one predicts a set of labels to associate with a given object. For example, an article can belong to multiple categories; an image can be associated with several tags; in medical billing, a patient report is annotated with multiple diagnosis codes. The most commonly used approach, called binary relevance (BR), trains one binary classifier to predict each label separately. BR ignores label dependencies and often makes conflicting predictions, such as tagging "cat" but not "animal" for an image. How to learn label dependencies from data and train classifiers to account for such dependencies is the central question in multi-label classification. This thesis describes two new approaches to multi-label classification which leverage label dependencies and achieve better classification accuracy than BR and many other sophisticated methods. The first approach, called conditional Bernoulli mixture (CBM), directly estimates the joint probability distribution among all labels with a mixture model. CBM's special model structure allows for efficient training, joint inference and marginal inference procedures designed to optimize different metrics. The second approach, named BR-rerank, seeks to improve both BR's confidence estimation and prediction through post calibration and reranking procedures. BR-rerank takes the BR predicted set of labels and its product score as features, extracts more features from the prediction itself to capture label constraints, and applies Gradient Boosted Trees (GB) as a calibrator to map these features into a calibrated confidence score. The GB calibrator not only produces well-calibrated scores (aligned with accuracy and sharp), but also models label interactions, correcting a critical flaw in BR. Experiment results show that reranking label sets by the new calibrated confidence makes accurate set predictions on par with state-of-the-art multi-label classifiers-yet calibrated, simpler, and faster. Both CBM and BR-rerank are reduction methods: they transform a complex multi-label classification problem to a series of standard binary classification, multi-class classification and regression problems, which are easier to solve.