Engineering Text-to-text Generation Language Models as Discriminative Classifiers for Accurate Answer Detection

Mohammed Azam Sayeed, Deepa Gupta, Vani Kanjirangat · Procedia Computer Science · 2025

Online peer grading is gaining wide adoption to manage open-ended questions involving students in the grading process. Evaluators could save a lot of grading time by using semi-automating peer grading to identify the most accurate/highest-scored and use them as anonymous graders promoting competency and fairness, thereby empowering students in the grading process. Leveraging language models (LMs) of smaller-scale large language models (LLMs), medium language models (MLMs) and small language models (SLMs) can offer greener, affordable AI (Artificial Intelligence) inferences for domain adaption suitable for limited educational settings. This study utilizes a university-sourced dataset compiled from advanced university courses like AI and ML (Machine learning), which are typically scarce in volume as seen in most education settings. We explore the efficacy of domain-adapting text-to-text generation language models (such as mT5, blenderbot, Flan-T5, byT5, and switch-base collections) with greener parameters to behave as discriminative classifier models via full finetuning, without changing the base model architecture having all base parameters as trainable, unlike partial base model finetuning with adapters. Our research establishes the flan-T5-large model to achieve about 83% for both accuracy and f1 score. These experimental models are analyzed from various perspectives such as topics, questions, and marking schemes on the standard evaluation metrics.

Read the paper · More papers on PaperTik