Corpus Variations for Translation Lexicon Induction

Rebecca Hwa, C. Bailey Nichols, Vivisimo Inc · UvA-DARE (University of Amsterdam) · 2006

Lexical mappings (word translations) between languages are an invaluable resource for mul-tilingual processing. While the problem of extracting lexical mappings from parallel cor-pora is well-studied, the task is more challeng-ing when the language samples are from non-parallel corpora. The goal of this work is to investigate one such scenario: finding lexical mappings between dialects of a diglossic lan-guage, in which people conduct their written communications in a prestigious formal dialect, but they communicate verbally in a colloquial dialect. Because the two dialects serve dif-ferent socio-linguistic functions, parallel cor-pora do not naturally exist between them. An example of a diglossic dialect pair is Modern Standard Arabic (MSA) and Levantine Ara-bic. In this paper, we evaluate the applicabil-ity of a standard algorithm for inducing lexical mappings between comparable corpora (Rapp, 1999) to such diglossic corpora pairs. The fo-cus of the paper is an in-depth error analysis, exploring the notion of relatedness in diglossic corpora and scrutinizing the effects of various dimensions of relatedness (such as mode, topic, style, and statistics) on the quality of the result-ing translation lexicon. 1

Read the paper · More papers on PaperTik