Hidden Cost of Mutation Testing on Auto‐Grader

Rifat Sabbir Mansur, Clifford A. Shaffer, Stephen H. Edwards · Computer Applications in Engineering Education · 2025

ABSTRACT Mutation testing (MT) is a powerful technique for evaluating the quality of software test suites. MT introduces faults or “mutations” into the code and checks whether the tests then fail as appropriate. While MT is known to be more effective than code coverage as a measure of test quality, its computational cost makes it challenging to deploy in educational settings. In this paper, we show the effects of this computational demand on an auto‐grading system when MT was used in a junior‐level Data Structures and Algorithms (DSA) course. Through a comparative study spanning semesters with and without MT, we observed a noticeable increase on the auto‐grader's processing time and feedback turnaround time (about 30–50 s, which represents roughly a tripling in per‐submission processing time) for students whose projects are graded with MT. This additional load raises concerns that it might overload the server, causing delays for students in other courses. However, with suitable mitigation strategies in place, the only measurable impact on other students was a higher variance in feedback turnaround times during peak use. One such mitigation strategy is the use of a local MT plug‐in which helped to reduce the total number of submissions to the auto‐grader. Overall, we find the effects on server load from a carefully chosen set of mutations combined with moderate use of local MT to have an acceptable computational cost on the system load while improving student test suite quality.

Read the paper · More papers on PaperTik