Classify Text Using K-Nearest Neighbor Algorithm to Reduce the Term Vector Space

Sandeep Kumar Sunori, M Chethan, Vinay Gaikwad, G. Ravivarman, Preetjot Singh, P. Sharmila · 2024

An essential part of text mining is text categorization, which involves the automatic sorting of many documents. The area of text categorization has benefited greatly from the introduction of several supervised learning algorithms. One of the greatest text classifiers because to its simplicity and effectiveness is the K-Nearest Neighbor (KNN) algorithm, which is one of these helpful strategies. One common approach to text classification is the vector space model, which involves representing a single document as a vector with a sequence of chosen words called feature items. A vector space model-based algorithm is KNN. Nonetheless, according to the conventional KNN algorithm, every feature item's weight in different categories is the same. I don't think this makes any sense. Because these features may have varying degrees of relevance and distribution across several categories. We propose a new and improved KNN method that takes variance into account, in light of this drawback of the original KNN algorithm. The results of the experiments demonstrate that the refined upgraded KNN outperforms the conventional KNN classifier.

Read the paper · More papers on PaperTik