Comparative Analysis of Encoder-Based Pretrained Models: Investigating the Performance of BERT Variants in Indonesian Question-Answering

Javier Islamey, Vannes Jonathan, Muhammad Nurzaki, Henry Lucky · 2024

Question-Answering (QA) is an NLP task designed to answer questions automatically. In Indonesian QA systems, various methods, such as rule-based, semantic-based, and machine-learning approaches with deep learning, have been implemented. The unique linguistic features of Bahasa Indonesia may pose significant challenges for some of these methods. Additionally, research on generative QA systems using pre-trained models in Bahasa Indonesia is underexplored compared to English. To address these issues, this paper aims to develop a deep learning-based question-answering system for Bahasa Indonesia using a sequence-to-sequence approach with pre-trained Bidirectional Encoder Representations from Transformers (BERT) models. The proposed solution implements IndoBERT, IndoRoBERTa, and mBERT in an encoder-decoder architecture with each BERT model as both the encoder and the decoder. The models employ different weight configurations as part of an ablation experiment. The models are then evaluated on the SQuAD-ID dataset. Results have shown that IndoBERT2BERT with pre-trained weights in both encoder and decoder have consistently outperformed every model on BLEU, ROUGE-N, BERTScore, EM, and F1. This proves IndoBERT2BERT's capability to generate texts that are lexically and semantically similar to the target answer while capturing relevant information. Findings also show that IndoBERT2BERT is particularly capable of handling longer answers, showcasing its ability to work with complex passages.

Read the paper · More papers on PaperTik