Essay Scoring with LLM Agent

Fernando Jonathan Polanco Espino, Anton Yu. Dolganov · 2025

This study presents an AI-based agent designed to automate the grading of student essays in the Ecology course at Ural Federal University. The agent leverages the capabilities of Large Language Models (LLMs), particularly from the Llama family, to evaluate various aspects of student writing, including grammar, structure, and content relevance. The system integrates advanced natural language processing techniques, Python libraries for data handling and analysis, and prompt engineering strategies to ensure accurate interpretation of student input. In addition to scoring, the model generates detailed feedback aimed at helping students understand their strengths and areas for improvement.Preliminary results from experiments using different versions of Llama models show a tendency to assign mid-range scores across most submissions. This indicates a challenge in clearly distinguishing between low- and high-quality responses, particularly for essays lacking strong arguments or cohesive structure. Despite occasional inconsistencies in scoring, the implementation of AI in the grading process has the potential to significantly streamline assessment workflows, reduce instructor workload, and enhance the consistency and depth of feedback. These findings suggest promising directions for improving both the accuracy and transparency of automated grading systems in educational contexts.

Read the paper · More papers on PaperTik