Evaluating Back Translation and Misspelling Correction Utilization on Indonesian AES

Elvina Amadea Tanaka, Steven Christian, Anderies, Andry Chowanda · 2024

This paper aims to tackle the problem of Automatic Essay Scoring (AES), a method used to predict whether an answer is correct given a question based on the provided guidelines and criteria. AES helps teachers to score student answers based on the guidelines fed into the model, saving time and decreases the human-error rate. We had chosen the UKARA challenge dataset, an Indonesian binary classification dataset consisting of two different questions separated into two datasets A and B. Prior studies utilized a stacking approach with XGBoost and a Neural Network model, achieving F1-scores of 88.4% and 75.9% for problem A and B respectively. The latest study utilizing the UKARA dataset used SBERT sentence embeddings and a Neural Network model, resulting in F1-scores of 89.4% for problem A and 75.7% for problem B. Although the F1-score for problem A was satisfactory, the F1-score for problem B remained relatively low. Therefore, this research aim to improve the performance of problem B. This research also discusses the various data augmentation techniques that can be applied to the dataset, such as back translation and misspelling correction Peter Norvig method. Based on the latest research, we decided to use SBERT sentence embeddings and neural networks for the model. The experiment managed to improve the result for problem B with a maximum F1-score of 77.2% along with 89.7% for problem A.

Read the paper · More papers on PaperTik