The IIT Bombay Hindi-English Translation System at WMT 2014

Piyush Dungarwal, Rajen Chatterjee, Abhijit K. Mishra, Anoop Kunchukuttan, Ritesh Shah, Pushpak Bhattacharyya · 2014

In this paper, we describe our English-Hindi and Hindi-English statistical systems submitted to the WMT14 shared task.The core components of our translation systems are phrase based (Hindi-English) and factored (English-Hindi) SMT systems.We show that the use of number, case and Tree Adjoining Grammar information as factors helps to improve English-Hindi translation, primarily by generating morphological inflections correctly.We show improvements to the translation systems using pre-procesing and post-processing components.To overcome the structural divergence between English and Hindi, we preorder the source side sentence to conform to the target language word order.Since parallel corpus is limited, many words are not translated.We translate out-of-vocabulary words and transliterate named entities in a post-processing stage.We also investigate ranking of translations from multiple systems to select the best translation.

Read the paper · More papers on PaperTik