Grammar Error Correction in Indonesian Sports News: Comparing the Performance of Pre-Trained T5 and BART Models

Muh Adnan, Esa Prakasa, Yuyun Yuyun, Najirah Umar, Nasrullah Nasrullah, Adi Sadli, A.Edeth Fuari Anatasya, Hazriani, Mashur Razak · 2024

This research compares the performance of two pre-trained models based on the transformer, namely text-to-text transfer transformer (T5) and bidirectional and auto-regressive transformer (BART), through a fine-tuned for correcting grammar in Indonesian sports news texts. We used vennify/t5-base-grammar-correction and onionLad/grammar-correction-bart-base, was trained using a dataset of 10,075 sentences from Indonesian language sports news sites. We tested the model performance using bilingual evaluation understudy (BLEU), google's language evaluation understanding (GLEU), training loss, and validation loss. The results show that BART was more efficient in learning data patterns and has better generalization ability, supported by BLEU and GLEU scores of 0.8560 and 0.9655. Meanwhile, T5 maintains simpler sentence structures, supported by BLEU and GLEU scores of 0.8120 and 0.9564. BART excels at handling more complex contexts, making it more effective at accurately correcting errors in dynamic sports news. Both models have their respective advantages, with BART being better at correcting detailed grammatical errors, while T5 offers consistent correction of more simple sentence structures.

Read the paper · More papers on PaperTik