Estimating Word Translation Probabilities from Unrelated Monolingual Corpora Using the EM Algorithm

Philipp Koehn, Kevin K. Knight · 2000

Selecting the right word translation among several options in the lexicon is a core problem for machine translation. We present a novel approach to this problem that can be trained using only unrelated monolingual corpora and a lexicon. By estimating word translation probabilities using the EM algorithm, we extend upon target language modeling. We construct a word translation model for 3830 German and 6147 English noun tokens, with very promising results. 1. Introduction Selecting the right word translation among several options in the lexicon is a core problem for machine translation. The problem is related to word sense disambiguation, which tries to determine the correct sense for a word occurrence (e.g. river bank vs. money bank). While the definition of word sense is a tricky issue, the picture is much clearer in translation. If we observe human translators, we can collect up the different ways in which a German word is usually translated into English. In some contexts, ...

Read the paper · More papers on PaperTik