Training Data Augmentation for Code-Mixed Translation

Abhirut Gupta, Aditya Vavre, Sunita Sarawagi · 2021

Machine translation of user-generated codemixed inputs to English is of crucial importance in applications like web search and targeted advertising.We address the scarcity of parallel training data for training such models by designing a strategy of converting existing non-code-mixed parallel data sources to codemixed parallel data.We present an mBERT based procedure whose core learnable component is a ternary sequence labeling model, that can be trained with a limited code-mixed corpus alone.We show a 5.8 point increase in BLEU on heavily code-mixed sentences by training a translation model using our data augmentation strategy on an Hindi-English codemixed translation task.

Read the paper · More papers on PaperTik