K-Nearest Neighbors in Machine Learning: Theory, Practice, and Advances
N. Manjunathan, R. Vijaya Prabhu, K. Sudha, D. Lekha, M. Ezhilvendan, Gayathri S · 2024
K-Nearest Neighbors (KNN) is a basic model in a ML field used for classification or prediction analysis owing to its efficiency. The following paper will be a survey paper focused on elaborating the theory of KNN, its distance learning nature and effects of the selection of (k) and distance measure. We consider different options of KNN and the generalization of the algorithm such as weighted KNN and other distance measurement methods and note the feature scaling. Subsequently, we look at the use of KNN in different domain including; healthcare, finance and investing, image recognition and recommender systems to note its versatility. This paper also reviews the strengths and weaknesses of KNN: this technique belongs to the class of non-parametric methods, but it may turn out to be sensitive to data with high dimensionality, and it is also known to be slow as a method when it works with big data. In an effort to improve KNN, modern trends in dimensionality reduction, ensemble learning, and distributed systems are discussed and how these advancements overcome the old problems. Last, the paper summary discusses how KNN is still used in present day machine learning and how further enhancements could be made. Thus, this paper's purpose is to offer an extended description of KNN so that it will enable users to apply the method proficiently in different practical problems.