K-Nearest Neighbour Classifier for Big Data Mining based on Informative Instances
Proma Hossain Progga, Md. Jobayer Rahman, Swapnil Biswas, Shakil Ahmed, Dewan Md. Farid · 2023
K-Nearest Neighbour (KNN) a well-known method that is using in several real-life applications that involve machine learning and data mining in the real world. It finds the nearest neighbors from the training data and classifies the test data using majority voting. However, KNN does not build a model or decision line based on training data, and it can become computationally expensive with large datasets. To address this issue, our proposed strategy divides the Big Data into clusters and extracts informative instances from each cluster based on the cluster center, the closest neighbors of the center, and randomly selected instances. The informative instances are then used for KNN classification, resulting in significant savings in time and storage space without compromising the classification performance. Experimental results demonstrate that our proposed strategy improves the effectiveness of the clustering and KNN algorithms, making it suitable for real-world applications with large-scale datasets.