HYDBre: A Hybrid Retrieval Method for Detecting Duplicate Software Bug Reports

JiaWei Zhang, Lei Xiao, Meiguang Li, Zhaorui Meng, YiHao Li · 2024

With the continuous growth in the size and complexity of software systems, the number of software defects has rapidly increased, making bug reporting a critical aspect of software development and maintenance. However, inconsistent descriptions often lead to duplicate bug reports, which in turn increase developer workload and waste resources. Existing methods for detecting duplicate bug reports (DBRD) are typically framed as either information retrieval (IR) tasks or classification tasks. In this paper, we approach the DBRD task as a retrieval problem. Traditional IR-based methods, while effective at retrieving information, often lack the ability to capture deep semantic relationships. On the other hand, deep learning methods excel at understanding semantics but require large amounts of training data to perform well. This poses a challenge for smaller projects with fewer duplicate reports, where training models from scratch becomes impractical due to limited data. To overcome this limitation, our approach focuses on fine-tuning pre-trained models, reducing the dependence on large datasets. To address these challenges, we propose HYDBre, a hybrid method that integrates traditional DBRD retrieval techniques with sentence embedding-based retrieval. Specifically, we fine-tune the BGE-M3 model and combine the retrieval results using a preferential algorithm. We evaluate HYDBre on three industrial datasets and benchmark it against four existing methods. Experimental results demonstrate that HYDBre achieves state-of-the-art performance, with Recall Rate@10 scores ranging from 0.704 to 0.794 across all datasets. Our method outperforms previous state-of-the-art approaches by 16.94% to 24.65% in terms of recall, validating its effectiveness in solving the DBRD task.

Read the paper · More papers on PaperTik