Study on KNN arithmetic based on cluster
Yang Bing-ru · Jisuanji gongcheng yu sheji · 2009
Traditional KNN arithmetic compares with every sample vector in sample space in order to find k neighbors of classification of the sample.This causes computing times too much and system performance degrades.So,the traditional KNN arithmetic,clusters training document with highly overlapping word is improved,central vector of cluster is gained.In the text classification process,first comparability is compared with central vector of each cluster,then comparability is compared with each document in cluster when comparability with central vector reach threshold.Computing times are reduced at a certain extent.At the same time,improve the IF-IDF formula so as to term’s position in the text is different,it should have difference weigh.