Improving Performance of Automatic Duplicate Bug Reports Detection using Longest Common Sequence : Introducing New Textual Features for Textual Similarity Detection

Behzad Soleimani Neysiani, Seyed Morteza Babamir · 2019 5th Conference on Knowledge Based Engineering and Innovation (KBEI) · 2019

Automatic duplicate bug reports detection is a famous problem in mining software repositories since 2004 for software triage systems e.g. Bugzilla. Textual features are the most important type of features in similarity and duplicate detection e.g. BM25F which indicate the common term frequency in two reports. Sometimes a common sequence can show more similarity in two texts, thus new features based on longest common sequence (LCS) of two bug reports proposed in this paper as new textual features for text similarity detection. Android, Eclipse, Mozilla, and Open Office dataset are used for evaluation of proposed features and the experimental results show LCS-based features are important and the accuracy, precision and recall of classifier prediction models improved 4.5, 2.5 and 2.5 percent respectively on average after using LCS and get up to 96, 98 and 97 percent respectively on average using different classifiers.

Read the paper · More papers on PaperTik