Domain Adaptation for Hindi-Telugu Machine Translation using Domain Specific Back Translation
LTRC, IIIT Hyderabad, India, Hema Ala, Vandan Mujadia, LTRC, IIIT Hyderabad, India, Dipti Misra Sharma, LTRC, IIIT Hyderabad, India · 2021
In this paper, we present a novel approach for domain adaptation in Neural Machine Translation which aims to improve the translation quality over a new domain.Adapting new domains is a highly challenging task for Neural Machine Translation on limited data, it becomes even more difficult for technical domains such as Chemistry and Artificial Intelligence due to specific terminology, etc.We propose Domain Specific Back Translation method which uses available monolingual data and generates synthetic data in a different way.This approach uses Out Of Domain words.The approach is very generic and can be applied to any language pair for any domain.We conduct our experiments on Chemistry and Artificial Intelligence domains for Hindi and Telugu in both directions.It has been observed that the usage of synthetic data created by the proposed algorithm improves the BLEU scores significantly.