Wikipedia Based Semantic Related Chinese Words Exploring and Relatedness Computing

Yixin Zhong · Beijing Youdian Xueyuan xuebao · 2009

To find how to collect semantic related words and calculate semantic relatedness,an experiment is done to download about 50 thousand documents from the web site of Chinese Wikipedia and extract hyperlinks between lines which contains semantic information.By mining hyperlinked references in documents,about 400 thousand semantic related word pairs are collected.With more experiments on topic groups of related words,tightly related words are grouped into smaller sets with an average semantic relatedness calculated.Semantic relatedness is calculated using information of hyperlink positions and frequencies in documents.Comparing with the result by classic algorithms,the reliability of the new measures is analyzed.

Read the paper · More papers on PaperTik