Text similarity calculation based on language network and semantic information
Zhijia Zhan · Computer Engineering and Applications Journal · 2014
Aiming at the shotcoming of traditional text similarity methods with statistical information of word frequency and semantic information of word in text, it proposes a new text similarity calculation based on language network and word semantic information. This new method extracts feature items based on the feature values of the word nodes in a documental language network. It also considers both the importance of feaure items and the semantic relations among feature items, and proposes to construct a semantic network of document feature items to calculate the similarity of documents. Finally it uses several K-means clustering methods for evaluating preformance of the new text document similarity. Experimental results show that the method's F-measure is superior to the others' which proves that the proposed method is effictive.