Comparative Study of LLMs in Assessment Evaluation and Affective Feedback Generation
Sreenandana Nandakumar, Sidharth S Kumar, Sruthy Anand, Sethuraman N Rao · 2025
Artificial Intelligence in education has seen a great surge recently, especially in transforming traditional education systems as students use tools like ChatGPT for their daily learning especially in higher education. AI and LLM have great potential in transforming even the traditional assessment methods by automating time-consuming processes, such as manual assessment, quiz evaluation, feedback generation, and grading. In order to build a system capable of assessment and student evaluation using LLM, it is necessary to evaluate LLMs capability in finding mistakes in the student answer and to provide feedback. To achieve this, our study aims to evaluate the LLMs and select the best LLM based on its ability to identify the mistakes and provide descriptive feedback. In this paper we present a comparative analysis of four popular LLMs, namely ChatGPT, Claude, Gemini, and MetaAI, in scoring and feedback generation for a given reference material and set of questions and answers. The models are assessed based on their scoring accuracy and the quality of the generated feedback. The study finds that LLMs are excellent in finding mistakes and giving accurate feedback for improvement. However, our test result shows GPT4O provided the most qualitative overall feedback when compared to the other LLMs, especially in terms of specificity and the constructive and suggestive phrases used. The paper also assessed these LLM feedback based on sentiment polarity, actionable suggestion and specificity in feedback.