The Research of Chinese Semantic Similarity Calculation Introduced Punctuations

Cheng Xianyi -, Sun Ping -, Qian Zhu, Cai Yue-hong · Journal of Convergence Information Technology · 2010

So far, most Chinese natural language processing neglects the punctuations or oversimplifies their functions. To improve the efficiency of Chinese similarity computing, this paper gives a Chinese similarity computing system model in accordance with the problems of Chinese sentence similarity computation aspect. This model is a combination of punctuations and traditional similarity computing. Comparing with Cosine-based Similarity calculation, the sentence similarity calculation based on word shape and word order, and Sentence similarity calculation based on semantic, this model makes up their insufficiency to a certain extent. The experiment shows that this model has a higher rate of accuracy in Chinese similarity computing.

Read the paper · More papers on PaperTik