Using SVM for classification in datasets with ambiguous data
Seyyed Mohammad Reza Hashemi, Thomas P. Trappenberg · 2002
One of the challenges in machine learning is the classification of datasets with ambiguous instances. In this paper we study specifically datasets with examples that have overlapping feature values for different classes. In these circumstances there is a bound on the classification performance. While there seems to be a race for accuracy, very little has been done to understand and solve the issues related to ambiguous data where the possible classification performance is limited. We discuss the use of SVMs in a proposed scheme to handle classification in such problem domains. A new approach is offered that tries to separate ambiguous data from the data that are much simpler to classify in order to prevent their influence on the classification process. We demonstrate that by separating the ambiguous data, although we lose some data, the performance of the classification increases significantly. In contrast to previous findings with some other classifiers, our experimental results show that the performance of SVM classifiers on cleaned data is not affected significantly when there are some atypical points in the training data.