Euclidean distance based label noise cleaning
Muhammad Ammar Malik, Moonsoo Kang · 2017
Quality of the datasets play an important role in performance of supervised classifiers. In presence of mislabelled examples, the performance of such classifiers degrades severely. In this paper we propose a label noise cleaning approach based on euclidean distance. The examples most likely to be mislabelled generally have similar mean euclidean distances with positive and negative examples. Selecting such examples for an expert's review can help in cleaning the label noise. Since support vector examples contain most of the mislabelled examples, we propose that reviewing the top 50% support vector examples with similar euclidean distances for positive and negative examples can clean most of the label noise from the datasets. An important contribution of our proposed method is that it requires only 1 iteration to clean the label noise, which can be helpful in cases where datasets need to be made quickly. We show that our proposed method removed more than 90% of the label noise with relatively less examples reviewed.