TLFL: Token-Level Fault Localization for Novice Programs via Graph Representation Learning
Liu Yong, Ruishi Huang, Jizhe Yang, Binbin Yang, ShuMei Wu · 2024
Fault Localization (FL) is a potential method to enhance coding productivity for novices. However, current work on FL concentrates on method-level or statement-level. Despite hints of potential fault locations, it is still challenging to identify the root causes (i.e., faulty tokens) for novices, who have little familiarity with the programming language. In our research, we introduce TLFL, an innovative approach to Token-Level Fault Localization tailored for beginners. This method leverages comprehensive code architecture and coverage data, employing graph-based representation learning to enhance effectiveness. Specifically, we introduce a novel graph-based token-level representation designed to preserve code structures from various granularities, along with coverage information at the statement level. Then, we utilize a Graph Neural Network to derive significant features from the graph-based token-level representation and identify faulty tokens in the program. Evaluation results on 2189 student programs from the well-established standard benchmark CoderFlaws show that TLFL localizes 94 faults in Top-1, outperforming the state-of-the-art FL techniques, including RNNFL, CNNFL, SBFL, and FLITSR. In particular, TLFL localizes more 93, 166, 176, and 133 faults than the best-performing technique RNNFL in Top-1, Top-3, Top-5, and Top-10.