Impact of corpus size and quality on English-Bangla statistical Machine Translation system
Ali Hasan Imam, Miah Raihan Mahmud Arman, Shahadat Hossain Chowdhury, Khalid Mahmood Malik · 2011
Statistical machine translation (SMT) evolves with the motivation of translating a text from source language to target language which employs the machine learning technique to a parallel corpus for producing a translation system exclusively automatic. We have developed Anubad[26], a phrase-based Bangla to English SMT on the top of the SMT model proposed in [1] which is publicly available on www.anubad.com. As the most challenging task for SMT system development is the designing of large parallel corpora as the translation quality significantly depends upon the corpus dimension and quality, Bangla parallel corpus suffers the same problem and fails to provide a standard translation till now. In this paper, through simulations, we provide a guideline for developing an English-Bangla bilingual corpus. Although in a phrase-based Statistical Machine Translation systems, more training data is generally better outcome, however, we deflect from this notion and according to our experimental results, we observed that quality of good corpus could significantly improve the Bangla to English translation quality. We have found better translation quality by employing our techniques and achieved effective improvements on NIST and BLEU scores.