Classifiers Based on Inverted Distances
Marcel Jiřina, Marcel Jiřina · InTech eBooks · 2011
In this chapter we describe an elaborated yet simple classification method (IINC) that can outperform a range of standard classification methods of data mining, e.g. k-nearest neighbors, Naive Bayes Classifiers’ as well as SVM. In any case the method is an alternative to well-known and widely used classification methods. There is a lot of classification methods, simpler or very sophisticated. Some standard methods of probability density estimate for classification are based on the nearest neighbors method which uses ratio k/V, where k is the number of points of a given class from the training set in a suitable ball of volume V with center at point x (Silverman, 1990; Duda et al., 2000; Chaves et al., 2001) sometimes denoted as query point. For probability density estimation by the k-nearest-neighbor (k-NN) method in En, the best value of k must be carefully tuned to find optimal results. Often used role of thumb is that k equals to square root of number of samples of the learning set. Nearest neighbors methods exhibit sometimes surprisingly good results see e.g. (Merz, 2010; Kim & Ghahramani, 2006). Bayesian methods form the other class of most reputable non-parametric methods (Duda et al., 2000; Kim & Ghahramani, 2006). Random trees or random forest approach belong among complex, but the best classification methods as well as neural networks of different types (Bock, 2004). The disadvantage of many these methods is the necessity to find proper set of internal parameters of the system. This problem is often solved by the use of genetic optimization as in the case of complex neural networks see e.g. (Hakl et al., 2002). First, we will provide a shot overview of the basic idea of the IINC and its features and show a simple demonstative example of a pragmatic approach to a simple classification task. Second, we give a deeper mathematical insight into the method and finally we will demonstrate the power of the IINC on data sets from two well-known repository real-life tasks.