Text similarity calculation based on domain feature word
Yan Luo, Ouyang Ning · 2012
This paper proposes an improved method for feature selection based on traditional mutual information by establishing domain feature words which utilize the differences in the representation of a word in different classes. By the method, we can reselect the feature set out of the established one based on the traditional mutual information. It not only reduces the dimension of the vector but also represents the text more effectively. At the same time, a text similarity calculation system is designed in this paper. Finally, the experimental results show that the improved feature extraction method is superior to the traditional mutual information and the system has a good performance.