A Comparative Study of LSTM, Bi-LSTM, and BERT for Automated Essay Scoring
Khin Su Su Han, May Thu Myint · 2024
The rapid progress in Artificial Intelligence (AI) and natural language processing (NLP) has created new possibilities for educational technology, particularly in automated essay scoring (AES). This research explores the potential of Artificial Intelligence (AI) in revolutionizing essay scoring within educational technology. Specifically, it focuses on comparing the effectiveness of three distinct approaches to automated essay scoring (AES): training a Long Short-Term Memory (LSTM) model, Bidirectional Long Short-Term Memory (Bi-LSTM) from scratch, and fine-tuning a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model. Both models are trained and finetuned for each approach using the ASAP dataset from Kaggle’s Learning Agency Lab - Automated Essay Scoring 2.0 competition. Their performance is evaluated based on metrics like Mean Absolute Error (MAE), Mean Squared Error (MSE), and Accuracy. To facilitate practical application, a Python web application built with Streamlit enables users to submit essays, choose a scoring model (LSTM, Bi-LSTM or BERT), and receive feedback based on the selected model’s prediction. This study contributes valuable insights into the strengths and weaknesses of each approach, furthering our understanding of efficient model selection for AES tasks.