Automated Essay Assessment Using Generative AI: Evaluating DeepSeek’s Performance in University-Level Grading

Laxmisha Rai, Kejun Sheng, Fasheng Liu · 2025

The purpose of this paper to quantitatively evaluate the performance of DeepSeek while evaluating essay-style university student assessments. Different types of essays are collected, with varying degree of answers, features and are prompted to obtain their grade. To ensure a rigorous evaluation of ten essays considered in this study, we carefully designed the evaluation framework to yield meaningful outputs from DeepSeek. Our methodology incorporates assessments exhibiting diverse characteristics, including both AI-generated responses and authentic student submissions. Prompts are entered in different modes such as default DeepSeek mode, and DeepThink (R1) mode. The results obtained are compared with that of actual human instructor, and the analysis of results are provided. Study indicates that DeepSeek demonstrates efficiency in comprehensively analyzing essay content, critically assessing its quality, and generating scores that closely align with those assigned by human instructors. However, human oversight remains essential to ensure precise and definitive evaluation of final scores.

Read the paper · More papers on PaperTik