Improving SMT by Using Parallel Data of a Closely Related Language

Galu scaron ccaron aacute kov aacute Petra, Bojar Ond rcaron ej · Frontiers in artificial intelligence and applications · 2012

The amount of training data in statistical machine translation critically affects translation quality. In this paper, we demonstrate how to increase translation quality for one language pair by introducing parallel data from a closely related language. Specifically, we improve English→Slovak translation using a large Czech-English parallel corpus and a shallow MT system for Czech→Slovak translation. Several options are explored to identify the best possible configuration.

Read the paper · More papers on PaperTik