Semantic similarity and document feature bit-string for efficient K-nearest neighbor
Zhimin Chen · Jisuanji gongcheng yu sheji · 2013
To improve the efficiency of text categorization effectively,an improved k-nearest neighbor algorithm(KNN) based on semantic similarity is proposed.The semantics are combined and bit-string of features in the text.Taking into account the contribution of similar word in the text,the accuracy of classification and the recall rate are improved effectively.And the bit-calculation based on the bit-string of features in the text,can filter out similar text from training texts,which can overcome the problem of large computation of KNN.The analysis of algorithm and the experiments are used to prove that the computational efficiency is enhanced significantly and the accuracy of classification and the recall rate are also improved effectively.