Automated Grading of First-year HTML CSS using LLMs
Ms Jocelyne Smith · 2025
Rising student enrolments in introductory computer science courses at tertiary institutions are creating significant grading bottlenecks. Educators are struggling to manage increasing workloads, resulting in an inability to provide timely and personalised feedback, impairing effective student skill development. This issue is particularly pronounced in universities constrained by budgetary limitations. This study investigates the use of artificial intelligence (AI), specifically large language models (LLMs), for the automated grading of first-year HTML and CSS assignments. The first phase of the research focuses on developing an automated grading system. Advanced AI techniques, such as Prompt Engineering, Tool Use, and Multi-Agent Collaboration, are employed to improve feedback accuracy and relevance. The second phase compares the effectiveness of different LLMs. Local-inference LLMs (i.e. models running on the researcher's own machine, e.g., Llama, Gemma) are compared to API-inference models (models running on the cloud, e.g., ChatGPT). The models are compared in terms of processing time, computational cost, and accuracy in order to determine their suitability in providing personalised feedback. Our findings demonstrate that API-inference models, specifically ChatGPT, outperform others likely due to large-scale infrastructure, large diverse datasets and advanced reasoning capabilities. This research highlights ChatGPT-4o-mini as the most effective model for providing scalable and cost-effective grading solutions in educational settings. The model demonstrated a significant improvement in efficiency, being able to grade and provide detailed personalised feedback for 240 submissions in just 45 minutes. A task that would take a lecturer, who grades between 4 and 5 submissions per hour, 48 to 60 hours to complete. The research highlights AI's potential to significantly reduce grading workloads and improve educational outcomes, especially at institutions with limited budgets. This system could assist assessment grading at the UFS Department of Computer Science and Informatics, with future work focusing on refinement and expanding the approach to other coding languages such as C# and SQL.