A Comparison of Machine Learning and Neural Network Algorithms for an Automated Thai Essay Scoring
Suttichai Suriyasat, Sapa Chanyachatchawan, Nuengwong Tuaycharoen · 2023
Thai students have relatively low scores on the reading literacy assessment conducted by PISA. Various studies reported that reading skills could be improved by writing. However, essay scoring is a time-consuming task. An automated essay scoring system can support both teachers and students by reducing the teachers' workload and providing predicted scores as feedback to students. A number of recent studies have focused on automated essay scoring dataset that contains only essays written in English. Little to no research has been done on the automated essay scoring system for the Thai language. The aim of this study is to develop a Thai essay scoring system using machine learning and deep learning models that have been reported to achieve good performance. We also try to improve the performance of our models by adding essay attribute features. The models that were used in this study are logistic regression, kNN, SVM, random forest, gradient boosting, XGBoost, LSTM (bag-of-words), LSTM (w2v), BERT-based, and LSTM+CNN (BERT embedding). The models were evaluated by six metrics, including accuracy, Quadratic weighted kappa, precision, recall, and F1-score along with 10-fold cross-validation. The experimental results show that XGBoost outperforms other models considering the majority of best metric scores in each set. For deep learning models with automatically extracted features from the text, the LSTM with word2vec features model yielded better performance than other deep learning models.