Learning Software Bug Reports: A Systematic Literature Review
Guoming Long, Jingzhi Gong, Hui Fang, Tao Chen · ACM Transactions on Software Engineering and Methodology · 2025
The recent advancement of artificial intelligence, in particular Machine Learning (ML), has witnessed its significant growth in various software engineering research fields. Among them, bug report analysis is one of such examples as it aims to automatically understand, extract and correlate information from the reports with the help of ML approaches. Despite the importance of ML in automating and enhancing bug report analysis, a comprehensive review that systematically examines the state-of-the-art in this area is still lacking. In this article, we provide a systematic literature review on this promising research topic. Our review covers 1,825 papers, from which we extract 204 most relevant studies for detailed analysis. Based on the statistics and trends observed in these reviewed studies, we obtained seven key findings summarized as follows: (1) The extensive use of Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) and \( k \) -Nearest Neighbor ( \( k \) NN) for bug report analysis, noting the underutilization of more advanced models like BERT due to their complexity and computational demands. (2) Word2Vec and TF-IDF are the most common methods for feature representation, with a notable increase in deep learning-based methods in recent years. (3) Stop word removal is the most common preprocessing method, followed by tokenization and stemming. Structural methods surged post-2020. (4) Eclipse and Mozilla are the most frequently evaluated software projects, reflecting their prominence in the field. (5) Bug categorization is the most popular task, followed by bug localization, assignment, and severity/priority prediction, with a growing interest in bug report summarization driven by advancements in NLP. (6) Most studies focus on general bug types, but there is increasing attention on specific bugs such as non-functional and performance bugs. (7) Common evaluation metrics include F1-score, Recall, Precision, and Accuracy, but bug report related evaluation metrics have not received significant attention. The majority of studies prefer \( k \) -fold cross-validation for model evaluation. (8) Many studies lack robust statistical tests or effect size measurements. Finally, based on the key findings, we discover six promising future research directions, by which we hope, together with the findings, can offer useful insights to practitioners of this particular research direction.