Iterative learning of parallel lexicons and phrases from non-parallel corpora

Meiping Dong, Yang Liu, Huanbo Luan, Maosong Sun, Tatsuya Izuha, Dakun Zhang · 2015

While parallel corpora are an indispensable re-source for data-driven multilingual natural lan-guage processing tasks such as machine translation, they are limited in quantity, quality and coverage. As a result, learning translation models from non-parallel corpora has become increasingly important nowadays, especially for low-resource languages. In this work, we propose a joint model for itera-tively learning parallel lexicons and phrases from non-parallel corpora. The model is trained using a Viterbi EM algorithm that alternates between con-structing parallel phrases using lexicons and up-dating lexicons based on the constructed parallel phrases. Experiments on Chinese-English datasets show that our approach learns better parallel lexi-cons and phrases and improves translation perfor-mance significantly. 1

Read the paper · More papers on PaperTik