Effect of Distance Measures on K-Nearest Neighbour Classifier
Vandana Kalra, Indu Kashyap, Harmeet Kaur · 2022 Second International Conference on Computer Science, Engineering and Applications (ICCSEA) · 2022
Machine learning classifiers vary concerning their learning functions and approach. K–nearest neighbour (K-NN) is a versatile instance-based machine learning algorithm used for classification without building a learning model. Varying distance measures in K-NN for computing distance between instances affect the classification accuracy. Commonly used distance metrics are Euclidean, Manhattan, Minkowski and Mahalanobis. In this work, extensive experiments were carried out on diverse datasets. These datasets vary in terms of the domain, number and types of features. The classification accuracy results obtained by applying different distance measures in the K-NN classifier were compared and analyzed. The results show that the Mahalanobis distance performed better with more than 90% accuracy, than the Euclidean and other measures when applied on most datasets. The other two most significant inferences from this analysis are: first, the Mahalanobis distance considered the correlation between the features and the deviation of instances from the mean, and second, it is scale-invariant.