Collecting Comparable Corpora from the Web
Ahmet Aker, Evangelos Kanoulas, David Guthrie, Robert Gaizauskas · 2011
Statistical machine translation (SMT) relies on the availability of rich parallel corpora. However, in case of under-resourced languages, parallel corpora are not readily available. To overcome this problem previous work has recognized the potential of using comparable corpora as training data. A critical first problem with such an approach is actually identifying and gathering corpora with potential value in improving SMT systems. In this work we attempt to address this problem by exploiting News articles published on the Web to gather a large amount of comparable corpora for a number of different languages including English, German, Croatian, Estonian, Latvian, Romanian, Greek and Slovenian. Given a pair of comparable documents we extract phrases and employ human annotators to indicate whether these phrases are correct translations of each other (parallel units). We report the initial results of our evaluation over the English-German collected corpora. The results indicate that our gathering method leads to corpora that consist of good quality align-able phrases and thus they offer a large potential to improve machine translation.