Experimentation with NMT models on low resource Indic languages

Nikunj Bansal, Goutam Datta, Anupam Pratap Singh · 2021 Sixth International Conference on Image Information Processing (ICIIP) · 2021

Today’s Artificial Intelligence (AI) is data centric. Unlike earlier rule based systems where we used to write many rules to solve any specific problem, these days, we need to train our machine learning models with the help of huge corpus (data set). In this paper, we have discussed one of the important applications of AI and computational linguistic i.e. Machine Translation (MT) which translates one natural language to another automatically.MT industry has passed through different phases since its earlier popular approach such as statistical machine translation (SMT) systems and its other version such as phrase based SMT. Performance wise NMT always outperformed SMT on various aspects. However, this holds true only for the languages having large parallel corpora. For low-resource languages, it still remains suboptimal. In this paper, we have applied NMT to low resources Indian languages, i.e. English-Hindi. We used a basic LSTM based Seq2Seq model and an attention-based Seq2Seq model with fixed vocabulary size. We merged the corpus collected from various sources and preprocessed them for further use. We used the BLEU metric score for evaluation. We also evaluated the Google Translator to compare our experimental results with it.

Read the paper · More papers on PaperTik