Automated Essay Scoring for Japanese: The Effects of Continual Pre-Training on Domain-Specific Reference Data

Boago Okgetheng, Koichi Takeuchi · 2025

In recent automatic essay scoring task, various models based on pre-trained large language models such as BERT and GPT have been proposed, demonstrating improved QWK scores. The pre-training language models are trained on sources like Wikipedia and large amount of crawled texts. However, when applying these pre-trained models to the essay data in specific domains, it is expected that the models could achieve higher accuracy if the models had more domain-specific knowledge. In this paper, we examine the effects of augmenting pre-trained language models with domain-specific texts on the Japanese essay scoring task. We use the Open Calm GPT models, which are primarily trained on Japanese texts, and further train these models on domain-specific texts. Subsequently, we apply the pre-trained models to essay scoring tasks by fine-tuning them with scored essays corresponding to twelve different prompts. Experimental results demonstrate that in many cases, models trained on domain-specific texts during the continual pre-training exhibit higher QWK scores compared to those without continual pre-training. Furthermore, our findings indicate that higher QWK scores can be achieved when we apply the training method with updating weights across all layers.

Read the paper · More papers on PaperTik