Learning Bilingual Word Mappings
Pushpak Bhattacharyya · Machine Translation · 2015
This chapter focuses on learning bilingual word mappings from parallel corpora. Factor-based statistical machine translation was introduced in 2007, wherein lemma and morphological features are inserted on parallel corpora. It was shown that for a given level of accuracy, morphology brings down the parallel corpora requirement. Conversely, given a fixed amount of parallel corpora, the accuracy level goes up if morph analysis is applied. Translation model probabilities have to be computed from parallel corpora. The landscape of languages is complex, with unique requirements for translation of individual language pairs. Translation between such languages can be looked upon as taking place at the base of the Vauquois triangle. Since word alignment and word translation are mutually supportive, a principled algorithm can be used to find both. The translation task is almost over when words are translated and the translated words are positioned.