Resp-kNN: A probabilistic k-nearest neighbor classifier for sparsely labeled data
Adrian Calma, Tobias Reitmaier, Bernhard Sick · 2016
Over the past few years extensive research has been conducted to solve classification problems with help of machine learning techniques. However, machine learning is data-driven and obtaining labeled data is often challenging in real applications. Techniques that try to overcome this burden, especially, in the presence of sparsely labeled data, can be found in the field of semi-supervised or active learning, as both make use of unlabeled data. In this paper, a semi-supervised k-nearest neighbor classifier, called Resp-kNN, is proposed for sparsely labeled data. This classifier is based on a probabilistic mixture model and, therefore, combines the advantages of classifiers based on non-parametric density estimate (such as a classical k-nearest neighbor classifier based on Euclidean distance) and classifiers based on parametric density estimates (such as classifiers based on Gaussian mixtures). Experimental results on 21 publicly available benchmark data sets show that Resp-kNN is more robust (regarding the choice of k) and effective for sparsely labeled classification compared to several standard methods.