Atomic Evaluation using Large Language Model for Automated Essay Exam Scoring

I Made Widiana Putra Winarta, Syukron Abu Ishaq Alfarozi, Indriana Hidayah · 2024

Education is a cornerstone of national development and plays a crucial role in achieving sustainable development goals. Traditional manual grading of exams and assignments, particularly for subjective responses like essays, is time-consuming, resource-intensive, and prone to human biases and inconsistencies. To address these challenges, Automated Essay Scoring (AES) systems have been developed, leveraging advancements in machine learning, deep learning, and particularly in Large Language Models (LLMs). However, small parameter LLMs often struggle to maintain coherence over longer texts and accurately capture nuanced details, resulting in lower agreement rates with human raters. This study proposes a novel approach called Atomic Evaluation to enhance the performance of small parameter LLMs in AES tasks. Atomic Evaluation decomposes student responses into smaller, more manageable atomic units, enabling more granular and focused evaluations. By aligning these atomic units with detailed rubrics and using a multi-step scoring process, the method improves the precision and transparency of automated scoring. Experiments conducted on two datasets, eLOK and ASAP-SAS, using Llama 3 8B and Gemma 2 9B models, show that Atomic Evaluation significantly improves assessment accuracy, with an increase in Quadratic Weighted Kappa (QWK) values of up to $\mathbf{0 . 0 8}$ for Llama 3 8B and 0.23 for Gemma 2 9B. These results demonstrate the potential of Atomic Evaluation to provide a more reliable and efficient AES solution, especially in educational environments with limited computing resources.

Read the paper · More papers on PaperTik