BERT-based cross-project and cross-version software defect prediction
Binwen Sun · Applied and Computational Engineering · 2024
In recent years, deep learning-based software defect prediction has gained significant attention in software engineering research. This study aims to explore the application of the BERT model in the field of software defect detection. Traditional methods are constrained by manually designed rules and expert knowledge, which leads to limited accuracy and generalization ability. The strengths of deep learning methods lie in their capacity to capture complex semantic and contextual information in code. However, the effectiveness of deep learning models is hindered by the small scale of software defect datasets. To address this issue, we introduce BERT as a pre-trained model and construct a downstream task neural network, comprising a single-layer fully connected layer and a softmax classifier. Additionally, we evaluate four variants of BERT to enhance predictive performance. Through empirical studies on software defect prediction across different versions and projects, we find that utilizing the BERT pre-trained model significantly enhances predictive performance. The experimental results demonstrate that our model outperforms TextCNN by 8.99% in terms of AUC score and LSTM by 5.66%. In terms of the F1 score, our model surpasses TextCNN by 4.51% and LSTM by 15.57%. The primary contribution of this paper is the proposal of a cross-version and cross-project software defect prediction method, leveraging a lightweight BERT-based neural network. We also discuss the reasons for the observed variations in the performance of the four BERT variants during testing.