A Comparative Study of Model Variations: English-Nepali Language Pair

Basab Nath, Sagar Tamang, Shiladitya Munshi, Krishna Kant Pandey, Saroj Kumar, Princy Randhawa · 2024

With over 7,000 languages spoken worldwide, neural machine translation (NMT) is critical for enabling intercultural communication and global exchange. This research undertakes a comprehensive evaluation of gated recurrent unit (GRU) sequence-to-sequence models for English-to-Nepali translation. The comparative analysis systematically explores the impact of key training hyperparameters including epochs (0 to 30), GRU units (128 and 256), and batch size (64 and 128) on translation quality. The findings reveal peak translation performance at 10 epochs and a batch size of 128, achieving a BLEU score of 4.65 and demonstrating capable translation abilities. The runner-up 10 epoch, 128 unit model achieves close performance, indicating acceptable variations in units. Before reaching this peak efficiency, models severely underfit with inadequate training from $0-5$ epochs. Past 10 epochs, models drastically overfit with losses in generalizability and BLEU scores dropping to zero by 30 epochs.

Read the paper · More papers on PaperTik