Using $\mu \text{Bert}$ to Improve Mutation-Based Fault Localization for Novice Programs
Yuxing Liu, Jiaxi Li, Haiyang Wu, Xuchuan Zhou, Hengyuan Liu · 2024
Programming learning has emerged as a prevalent trend. However, due to their lack of experience, novices often encounter difficulties in effectively debugging programs. To address the challenges faced by novices in programming, the development of automatic program fault localization presents a potential solution. Nonetheless, the majority of existing fault localization methods prove inadequate for novice-level programs, with the two predominant approaches being mutation-based fault localization (MBFL) and spectrum-based fault localization(SBFL). The SBFL may falter, particularly with novice programs failing to clear any tests, which restricts its utility. Meanwhile, the MBFL doesn't fully account for the gap between simulated mutants and real faults, often covering just a fragment of actual issues. This results in reduced fault localization efficiency and higher operational costs, underscoring the challenges each method faces in various situations. To overcome this obstacle, this study proposes a novel fault localization method for novice programs, termed$\mu \mathbf{Bert}$Mutation-Based Novice Fault Localization$(\mu \mathbf{MB}- \mathbf{NFL})$. This method introduces a new type of mutant, which, by integrating contextual information, first generates a mask at the statement level within the source code, and subsequently produces a token for the segment of the mask most likely to simulate an actual failure. These mutants are referred to as Neural-mutants. Within this methodology, we incorporate the$\mu \mathbf{Bert}$tool, based on the CodeBERT model, to conduct mutation testing, substituting traditional mutants with the newly generated Neural-mutants for fault localization in novice programs. To evaluate the feasibility of the approach, experiments are conducted on 182 fault codes from novice programmers, using SBFL and MBFL as benchmarks. The experimental outcomes demonstrate that$\mu \mathbf{MB}- \mathbf{NFL}$surpasses the other two methods in terms of localization precision. Specifically,$\mu \mathbf{MB-NFL}$achieves localization for 51, 115, and 142 in the three indicators of$TOP - N (N=1,3,5)$, respectively.