Supervised Learning in Absence of Accurate Class Labels
Ramasubramanian Sundararajan, Hima Patel, Manisha Srivastava · Advances in computational intelligence and robotics book series · 2017
Traditionally supervised learning algorithms are built using labeled training data. Accurate labels are essential to guide the classifier towards an optimal separation between the classes. However, there are several real world scenarios where the class labels at an instance level may be unavailable or imprecise or difficult to obtain, or in situations where the problem is naturally posed as one of classifying instance groups. To tackle these challenges, we draw your attention towards Multi Instance Learning (MIL) algorithms where labels are available at a bag level rather than at an instance level. In this chapter, we motivate the need for MIL algorithms and describe an ensemble based method, wherein the members of the ensemble are lazy learning classifiers using the Citation Nearest Neighbour method. Diversity among the ensemble methods is achieved by optimizing their parameters using a multi-objective optimization method, with the objective being to maximize positive class accuracy and minimize false positive rate. We demonstrate results of the methodology on the standard Musk 1 dataset.