Transformer Hyperparameter Tuning for Madurese-Indonesian Machine Translation
Fika Hastarita Rachman, M Syauqi, Noor Ifada, Imamah Imamah, Sri Intan Wahyuni · Engineering Technology & Applied Science Research · 2025
The main problem arising in using Neural Machine Translation (NMT) for the Madurese language is the limitation of training data due to the unavailability of an adequate parallel corpus. In addition, the model must overcome the difference in words caused by the level of politeness in the Madurese language (coarse, moderate, and smooth). The rules-based approach requires many rules to represent these differences. In contrast, the statistical approach relies on the frequency of words in the training data, which cannot accurately capture variations in politeness levels. To overcome this problem, a parallel corpus was created to provide adequate training data, and an embedding matrix based on Skip Gram with Negative Sampling (SGNS) was used to produce better word representations for processing with transformers. This study also employs two types of evaluation: model configuration based on dataset size (large and small) and two tokenization methods (word and subword levels). The best results were obtained with the large dataset using word-level tokenization, achieving 0.70% accuracy for entirely correct text, 78.87% for partially correct text, and a BLEU score ranging from 4.76 to 27.63 with a maximum n-gram value from 1 to 4. This approach improved translation accuracy and shows significant potential for developing NMT systems for languages with limited resources, such as the Madurese language.