Analysis and Evaluation of Comparable Corpora for Under Resourced Areas of Machine Translation
Andrejs Vasiļjevs, Raivis Skadiņš, Robert Gaizauskas, Dan Ioan Tufis, Tatiana Gornostay · 2010
Lack of sufficient linguistic resources and parallel corpora for many languages and domains currently is one of the major obstacles to further advancement of automated translation. The solution proposed in this paper is to exploit the fact that non-parallel bi- or multilingual text resources are much more widely available than parallel translation data. This position paper presents previous research in this field and research plans of the ACCURAT project. Its goal is to find, analyze and evaluate novel methods that exploit comparable corpora in order to compensate for the shortage of linguistic resources, and ultimately to significantly improve MT quality for under-resourced languages and narrow domains.