Leveraging Prompt Engineering for Effective Automated Answersheet Evaluation
Raihana Rasaldeen, Ria Mariam Mathews, Stefi Marshal Fernandez, Irin Rose Jaison, Fabeela Ali Rawther · 2025
Manual grading of handwritten answer sheets is time-consuming and prone to human bias, while automated systems often suffer from OCR inaccuracies and limited contextual understanding, affecting the reliability of evaluation. This paper introduces an AI-driven framework that leverages prompt engineering and pre-trained language models (LLMs) to improve grading accuracy without requiring any model training. Text is extracted using Google Vision OCR and refined through carefully structured prompts, followed by logic-based evaluation of answers. A comparative analysis of raw and refined OCR outputs is conducted across three prominent OCR models-Tesseract, PaddleOCR, and Google Vision-to assess improvements in it's performance metrics. The results demonstrate how prompt engineering reduces OCR-induced noise, enhances structural clarity, and enables consistent, scalable evaluation. This paper highlights both the grading potential of large language models and the effectiveness of prompting, thereby establishing a reliable approach for AI-based academic assessment.