Performance Evaluation of English to Bodo Neural Machine Translation System with Varying Model Architecture and Vocabulary Size

Parvez Aziz Boruah, Shikhar Kr. Sarma, Kishore Kashyap, Simanta Kalita · 2023

This paper is about a work done on Neural Machine Translation of English-Bodo language pair using deep learning technique. Bodo is a language of northeastern part of India particularly in the state of Assam. The experiments are performed using English-Bodo parallel data collected from open sources, as well as created inhouse using expert linguists. The dataset was subjected to data filtering and cleaning. IndicNLP library and Mosesdecoder are used for tokenization for Bodo and English respectively. After tokenization, Byte Pair Encoding (BPE) technique was used for subword tokenization. Two transformer encoder and decoder models are built using different architecture. OpenNMT-py framework is used for building our models. Experiments have been performed on the two models using two different vocabularies of 8000 and 16000 sizes. Highest BLEU score of 10.43 was achieved on the testset from training Model 2 with 8000 vocabulary. The details of the results are analysed against different models.

Read the paper · More papers on PaperTik