Smoothing Class Frequencies for KNN Medical Article Classification
Kostas Fragos, Christos Skourlas · 2016
K-Nearest Neighbor (KNN) is one of the most popular algorithms for text classification. In many experiments researchers have found that the KNN algorithm accomplishes very good performance on different data sets. In a previous work [16], we proposed an algorithm, called lf-igf KNN, to classify medical articles using local from neighborhood and global from corpus class label frequencies to device a weighting scheme for ranking all data points in the training set. In this work, we modify this previous work going beyond simple counting, both smoothing class label frequencies and neighbor's distances. We provide by this way an alternative and more robust weighted scheme for KNN classification. The evaluation experiments on the collection of medical documents, called Ohsumed, show promising results and inspire us to use smoothing techniques to treat class occurrences in traditional KNN classification.