Pre-processing and Resource Modelling for English-Assamese NMT System

Mazida Akhtara Ahmed, Kishore Kashyap, Shikhar Kumar Sarma · 2023

Neural Machine Translation (NMT) modelling for low-resourced languages is usually under-explored lacking proper analysis in the pre-processing stages which could potentially influence the performance of the model. Assamese, an indigenous language of the North East India with over 16 million speakers, is no exception. This paper is an attempt towards exploration of the resources and tools necessary for developing a neural computational model for English-Assamese machine translation. Linguistic specialties are discussed. Resource identification and consolidation are highlighted. The paper highlights the technical intricacies associated with the language in the process of getting the input ready for NMT modelling. The various stages of preprocessing requirements to make the resources NMT ready are discussed. This includes normalization, tokenization, sub word tokenization and word embedding methods. With an open list of different NMT toolkits, this paper describes a roadmap for preparing Assamese machine translation.

Read the paper · More papers on PaperTik