Applying Large Language Models to Enhance the Assessment of Java Programming Assignments
Skyler Grandel, Douglas C. Schmidt, Kevin Leach · 2025
The assessment of programming assignments in computer science (CS) education traditionally relies on manual grading, which strives to provide comprehensive feedback on correctness, style, efficiency, and other software quality attributes. As class sizes increase, however, it is hard to provide detailed feedback consistently, especially when multiple assessors are required to handle a larger number of assignment submissions. Large Language Models (LLMs), such as ChatGPT, Claude, and Gemini, offer a promising alternative to help automate this assessment process in a consistent, scalable, and fair manner.