Bug Localization Based on Context-Aware Code Translation and Feature Fusion

Wenyuan Cheng, Ruixuan Zhang, Gaofeng Wang, Yuqian Chen · International Journal of Software Engineering and Knowledge Engineering · 2026

Bug localization refers to the process of identifying the locations of bugs in software through a series of testing and analysis methods during software development or maintenance. An important research direction focuses on localizing buggy components given a bug report. Existing bug localization methods can be broadly categorized into two approaches. The first is traditional information retrieval-based bug localization, which treats bug reports and source files as plain text and relies on keyword matching for localization. The second is deep learning-based bug localization, which leverages deep learning models to capture semantic information in bug reports and source files. Although deep learning models have advantages in semantic understanding, a significant semantic gap still exists between source files (written in programming languages) and bug reports (written in natural language), which impacts model performance. In this work, we propose a novel bug localization framework named WisdomLoc. This framework matches bug reports with the source files by extracting both information retrieval features and semantic features from the source files. Furthermore, WisdomLoc enhances the semantic representation of source files by incorporating natural language descriptions translated from the corresponding program instructions, using context-aware code translation techniques. This helps further bridge the semantic gap between source files and bug reports. Specifically, WisdomLoc consists of four modules. The first module extracts shallow semantic features from the source files based on their textual content and structural information. The second module captures deep semantic features by translating program instructions into natural language descriptions. The third module incorporates traditional IR-based software features. Finally, the fourth module integrates the outputs of the previous three modules to compute the final matching score between the bug reports and the source files. We evaluated the performance of WisdomLoc on four benchmark datasets. Experimental results demonstrate that WisdomLoc outperforms eight state-of-the-art models.

Read the paper · More papers on PaperTik