Tokenization effect on neural machine translation: an experimental investigation for English-Assamese

Mazida Akhtara Ahmed, Kishore Kashyap, Shikhar Kumar Sarma · 2023

Tokenization, as a research task, is mostly overlooked when dealing with machine translation as much emphasis is placed on modelling or data enhancement, not to speak for language pairs categorized as low-resourced. The current work takes up this task of an experimental analysis of tokenization on Neural Machine translation (NMT) on a resource-poor language pair: English-Assamese. Four implementations of tokenization available as library are selected for this study: Moses, Nltk, OpenNMT, IndicNLP. The efficiency of these tokenizers in handling the language specialities is also discussed. Both the source and target languages are tokenized with these methods and 12 different NMT models are developed by choosing different combinations of the tokenizers. Perturbations in the evaluation scores have been observed for each combination in both directions (En→As and As→En). It is striking to observe that tokenizing Assamese text with Moses produces the most degrading results. The tokenization effect is again confirmed by testing with the source text tokenized with different tokenizers wherein it was clear that the tokenization schemes followed during model training has to be adhered to produce the best results.

Read the paper · More papers on PaperTik