Boosting K-Nearest Neighbour (KNN) Classification using Clustering and AdaBoost Methods
Dewan Md. Farid, Nabila Sabrin Sworna, Ruhul Amin, Nazifa Sadia, M.K.M. Rahman, Nazmul Khan Liton, Md. Saddam Hossain Mukta, Swakkhar Shatabda · 2022 IEEE Region 10 Symposium (TENSYMP) · 2022
K-Nearest Neighbour (KNN) is a lazy learning algorithm which has been successfully practiced in several real-life machine learning applications. However, it has been titled as lazy learner because it is unable to learn a decision line from historical training data and uses the data itself for classification. KNN discovers the closest neighbors around test instances; At the same time considers the majority voting of the class label to classify the test instances. In the case of noisy datasets, the value of K should be higher. It is still struggling with the time complexity and better accuracy rate in some cases for noisy datasets. In this article, we proposed a technique to boost up the performance of KNN classification employing clustering and boosting techniques. The proposed method automatically determine the value of K based on the current situation of the neighbours instead of setting previous and random values. It also achieves significant accuracy and reduces the time complexity in compare with traditional KNN approach. The proposed method clusters the datasets and finds the majority voting of nearest neighbours from each cluster to classify the test data. We experimented the performance of suggested model on several real benchmark datasets from UCI and KEEL machine learning repositories and found that proposed model gives significant results which are relatively satisfactory in terms of accuracy, precision, f'1-score and Matthew's Correlation Coefficient (MCC).