Semantic Similarity Calculation of Chinese Word

Liqiang Pan, Pu Zhang, Anping Xiong · International Journal of Advanced Computer Science and Applications · 2014

This paper puts forward a two layers computing method to calculate semantic similarity of Chinese word. Firstly, using Latent Dirichlet Allocation (LDA) subject model to generate subject spatial domain. Then mapping word into topic space and forming topic distribution which is used to calculate semantic similarity of word(the first layer computing). Finally, using semantic dictionary "HowNet" to deeply excavate semantic similarity of word (the second layer computing). This method not only overcomes the problem that it’s not specific enough merely using LDA to calculate semantic similarity of word, but also solves the problems such as new words (haven’t been added in dictionary) and without considering specific context when calculating semantic similarity based on semantic dictionary "HowNet". By experimental comparison, this thesis proves feasibility,availability and advantages of the calculation method.

Read the paper · More papers on PaperTik