Word2Vec Duplicate Bug Records Identification Prediction Using Tensorflow

Hussain Mahfoodh, Mustafa Hammad · 2020

Bug duplication reporting is one of the most widespread software problems that cause inconvenience for the internal software stakeholders. It is useful for developers to eliminate redundant bug records where the fewer bugs duplicated records in bug reports documentation the more efficiently allocated resources are set to fix and enhance the software features. In this paper, the word embedding (Word2Vec) approach is used on four different software components from the Mozilla Core dataset with different sentence types through the duplicated bug category records to compare whether two given bug record descriptions are categorized as related bugs records. Besides, this paper proposes three different similarity measures and explores the accuracy of each measure. The study results show that the approach's accuracy is proportional to the existence of similar words within any of the two given two bug records descriptions. Additionally, we found that percentage of similarity accuracy is improved by finding the closest word using the Euclidean distance method than traversing for more index adjacent values within the trained word vector array.

Read the paper · More papers on PaperTik