Chinese Lexical Semantic Similarity Computing Based on Large-scale Corpus
Xueqiang Lv · Zhongwen xinxi xuebao · 2013
Automatic acquisition of similar words is one of the most crucial problems in natural language processing tasks,e.g.the query extension in information retrieval,pattern identification in machine translation,parser analysis and WSD.This paper focuses on Chinese semantic similarity computing based on large corpus,investigating the computation of context feature weight,the vector similarity measures,the window context vs.the dependency context,and the newspaper corpus vs.web corpus.Our experiments show that,in the web corpus,using window-based context combined with PMI weights function,the cosine measures gets the best semantic similarity results.