NMT for a Low Resource Language Bodo: Preprocessing and Resource Modelling

Simanta Kalita, Parvez Aziz Boruah, Kishore Kashyap, Shikhar Kumar Sarma · 2023

Implementation of neural machine translation for low-resource language requires special attention. In this paper we have discussed about the resource modelling and preprocessing steps for a low resource Indic language Bodo for the purpose of neural machine translation from English. To capture the essence of the Bodo language, syntactic and morphological characteristics of the language are discussed in this paper. It includes the steps of data preparation, tokenization and subword tokenization which are considered as the pre steps for preparing the training and test data. The research highlights multiple subword tokenizers to capture the rare words, unknown words and generate the vocabulary. It discusses the SentencePiece, Byte-pair Encoding (BPE) and WordPiece tokenizer for the Bodo language. We have shown the outcomes of these subword tokenizers with respect to Bodo language which will serve as a priori for the future researcher working in NMT for Bodo and any other low resource language in general. Proper sub-tokenization can aid in raising the caliber of neural machine translation that is the reason three different tokenizers are experimented with while preparing the dataset. This paper also discussed about Seq2Seq LSTM and Transformer Machine Translation architectures.

Read the paper · More papers on PaperTik