Bilingual Dictionary Extraction from Wikipedia

Kun Yu, Jun’ichi Tsujii · 2009

The way of mining comparable corpora and the strategy of dictionary extraction are two essential elements of bilingual dictionary ex-traction from comparable corpora. This paper first proposes a method, which uses the inter-language link in Wikipedia, to build compara-ble corpora. The large scale of Wikipedia en-sures the quantity of collected comparable corpora. Besides, because the inter-language link is created by article author, the quality of collected corpora can also be guaranteed. Af-ter that, this paper presents an approach, which combines context heterogeneity simi-larity and dependency heterogeneity similarity, to extract bilingual dictionary from the col-lected comparable corpora. Experimental re-sults show that because of combining the advantages of context heterogeneity similarity and dependency heterogeneity similarity ap-propriately, the proposed approach outper-forms both the two individual approaches. 1

Read the paper · More papers on PaperTik