Effect of Context and Tokenization on Machine Translation of Arabic Conversations on Social Media

Farzan Saeedi, Ghaniya Al Hinai, Khoula Al Kharusi, Abdulrahman Aal Abdulsalam · Procedia Computer Science · 2025

Machine translation of many Arabic dialects is still lagging behind Modern Standard Arabic (MSA). Prior research showed that different tokenization schemes affect the quality of translations produced by neural machine translation systems. In addition, the inclusion of context may improve the quality of the translation. In this study, we explore the effect of including different prior context window sizes and tokenization schemes on the final quality of machine translations of low-resource Omani Arabic dialects. We fine-tuned a transformer-based model OPUS-mt-ar-en from HuggingFace using a dataset of 876,689 Arabic words obtained from social media. Our experiments indicate that using larger prior context window significantly improves translation quality from 14.1 to 25.4 BLEU points. We also show that training a model with longer context window sizes gives continuous improvement in translation quality. Finally, the experiments indicate that performing morphological tokenization on Arabic text prior to translation by the neural model does not have impact on translation quality.

Read the paper · More papers on PaperTik