A Bidirectional Statistical Machine Translation System for Exploring the Performance of the Low Resource Language Pair English-Nepali
Amit Kumar Roy, Bipul Syam Purkayastha, Saptarshi Paul · 2024
The advances in Neural Machine Translation (NMT) systems have cast a shadow over Machine Translation (MT), which was previously dominated by Rule-based Machine Translation (RBMT) and Statistical Machine Translation (SMT). While Neural Machine Translation works effectively for languages with abundant resources, Statistical Machine Translation remains favored for low-resource languages such as Nepali. Nepali itself possesses distinctive linguistic features, characteristics, and scripts. This paper introduces a bidirectional Statistical Machine Translation (SMT) system for the Nepali-English language pair, which is considered low-resource. The open-source toolkit MOSES and a parallel text corpus with approximately 17 K sentences are both used by the system. To evaluate the system’s effectiveness, automatic evaluation metrics such as BLEU, F-Score, and METEOR were used. The system recorded scores of 21.13,53.32, and 38.29 for translating text from English into Nepali and 22.26, 57.52, and 27.81 for translating text from Nepali into English. Additionally, a comparison between translation performance with that of Google Translate, a standard neural network translation service, is done. When translating from English to Nepali, the proposed system outperforms Google Translate in terms of automatic evaluation scores, accuracy and precision.