Performance Comparison Study of Machine Learning, Embedding Vector, and Pre-trained Language Model for Automated Essay Scoring

Korean Society for Educational Evaluation, Jong-Im Park, Giljae Kim, Kangyun Park, Sook-ki Choi · 교육평가연구 · 2025

This study utilized scoring feature-based machine learning models(Extra Trees, Random Forest, LightGBM), embedding vector-based deep neural network models(OpenAI Embedding, Sentence-BERT, Universal Sentence Encoder), and pre-trained models(KLUE-RoBERTa-base, XLM-RoBERTa-base) to compare the performance of automated scoring models for Korean essay-type responses. Based on a dataset of 9,762 argumentative writing samples from middle and high school students, the models were trained and evaluated. The results showed that among the pre-trained models, XLM-RoBERTa-base achieved the highest accuracy among the embedding-based models, OpenAI Embedding demonstrated the highest accuracy and among the scoring feature-based models, Extra Trees performed the best with an accuracy. Furthermore, although pre-trained models showed the superior performance in the comparison, model selection should consider interpretability in conjunction with the purpose of automated scoring, and the development of hybrid models rather than single models is proposed for the future.

Read the paper · More papers on PaperTik