A Multilingual Dataset of Student Answers, Human Grading, and Multi-LLM Evaluations for Automated Assessment Research Using JorGPT

Jorge Cisneros-González, Natalia Gordo-Herrera, Iván Barcia-Santos, Yolanda Cerezo, Javier Sánchez-Soriano · Data · 2026

The increasing adoption of Large Language Models (LLMs) in higher education has created a need for high-quality, publicly available benchmarks for automated assessment. Existing datasets often rely on synthetic responses or lack detailed human feedback. This paper presents a multilingual dataset of 3041 authentic student answers to 50 open-ended Computer Science questions, collected from real university assessments during the 2025–2026 academic year. The dataset includes the original student responses (Spanish) and their parallel translations (English), instructor (or teacher) defined ideal answers, blind human grading with qualitative feedback, and structured evaluations from three state-of-the-art LLMs (DeepSeek-chat-V3.2, Qwen-flash-2025-07-28, Gemini-2.5-flash-lite-001) using a unified JSON schema. This resource enables reproducible research in automated grading, feedback generation, and cross-lingual educational NLP.

Read the paper · More papers on PaperTik