Chinese Lexical Semantic Similarity Computing Based on Large-scale Corpus

Xueqiang Lv · Zhongwen xinxi xuebao · 2013

Automatic acquisition of similar words is one of the most crucial problems in natural language processing tasks,e.g.the query extension in information retrieval,pattern identification in machine translation,parser analysis and WSD.This paper focuses on Chinese semantic similarity computing based on large corpus,investigating the computation of context feature weight,the vector similarity measures,the window context vs.the dependency context,and the newspaper corpus vs.web corpus.Our experiments show that,in the web corpus,using window-based context combined with PMI weights function,the cosine measures gets the best semantic similarity results.

Read the paper · More papers on PaperTik