Large Language Model-supported Intelligent Evaluation of Open-ended Tasks: Case Studies and Reflections Based on SOU

De Li Lin, Zhihui Wei, Ruifang Hope Sun, Qing Huo Liu · 2025

Developing execution-based proficiency requires the use of more open-ended tasks, which present significant assessment challenges due to the unstructured nature of their responses. These challenges include the absence of standard answers, difficulty in objectifying assessment indicators, subjectivity in teachers' evaluations, and the increased complexity of the assessment process. The development of Large Language Models (LLM) in artificial intelligence offers a potential solution to these issues. This study explores human-AI synergy in the assessment of open-ended tasks using LLM and aims to address the challenges of traditional assessment methods. The research involves a pilot implementation of intelligent assessment, focusing on innovative teaching scheme design, human-AI collaboration in assessment, result review and analysis, and iterative improvements. Findings indicate that LLM-assisted assessment can initially support the evaluation of open-ended tasks; however, reliability issues persist, necessitating manual review. While assessment criteria can be structured through rule-based guidance, the diverse nature of open-ended tasks demands greater adaptability in assessment frameworks. The study also found that using large models to assist teachers in evaluating students' open-ended tasks is feasible to a certain extent. However, there are risks of hallucinations in large models as well as ethical and moral concerns, which can be addressed through risk warnings and human-AI collaboration. Additionally, compared to assessments in other language-related writing tasks, setting intelligent evaluation metrics for open-ended tasks in management courses is relatively easier. At the same time, it places higher demands on teachers' professional application skills and requires strong support and patience from schools. As an exploratory study on AI applications in education, this research has practical implications but requires further investigation to enhance its implementation and impact.

Read the paper · More papers on PaperTik