Building a Web-Based Parallel Corpus and Filtering Out Machine-Translated Text
Alexandra Antonova, Alexey Misyurev · Meeting of the Association for Computational Linguistics · 2011
We describe a set of techniques that have been developed while collecting parallel texts for Russian-English language pair and building a corpus of parallel sentences for training a statistical machine translation system. We discuss issues of verifying potential parallel texts and filtering out automatically translated documents. Finally we evaluate the quality of the 1-million-sentence corpus which we believe may be a useful resource for machine translation research.