Automated Handwritten Essay Evaluation in Moodle: Leveraging Google Vision OCR and Mistral 7b
Charis Arlie Largo Baclayon, Kid Omar Rendon Costelo, Jeremy Jules Loyola Flores, Cherry Lyn C. Sta. Romana · 2024
Traditional grading methods are often time-consuming and subjective, increasing the difficulties in maintaining academic integrity against the background of easy access to online resources. Even as pre-written and AI-generated content become even more available, handwritten essays offer one of the most viable ways of stimulating real learning and original thought among students. This study assessed the accuracy and efficiency of AI-powered grading to improve education amid these challenges. More precisely, the paper investigated the possibility of automatically grading the handwritten open-ended reflection essays of students in the “Living in the IT Era” course by leveraging AI. Through Optical Character Recognition (OCR), together with a fine-tuned and trained Large Language Model (LLM) Mistral 7b, the system replicated human-grading decisions in evaluating essays comprehensively. The predicted score and human-graded score were compared in evaluating the system. Analysis using BERT Score revealed a high degree of correlation between the two, with consistent precision (88.41%), recall (83.33%), and F1-score (85.78%). External user testing also revealed a positive perception: the system is perceived as user-friendly (usability: 4.28), generally understandable (comprehensibility: 3.68), and with functionalities relevant to educators' needs (relevance: 4.0). This study adds to the expanding body of research on AI in education and lays the groundwork for future investigations in this area.