Performance Analysis of Deep Learning Models for Software Fault Prediction Using the BugHunter Dataset
Thị Minh Phương Hà · Journal on Information Technologies & Communications · 2025
Software fault prediction (SFP) involves the identification ofpotentially fault-prone modules before the testing phase in the softwaredevelopment lifecycle. By predicting faults early in the development process, the SFP process enables software developers to focus their efforts oncomponents that may contain faults, thereby enhancing the overall quality and reliability of the software. Machine learning and deep learningtechniques have been widely applied to train SFP models. However, theseapproaches face several challenges, including irrelevant or redundant features, imbalanced datasets, overfitting, and complex model structures.The NASA dataset from the PROMISE repository is the most commonlyused dataset for fault prediction. Recently, the BugHunter dataset withits substantially larger number of instances was explored to train the SFPmodels. In this study, we present the comparative study of three deeplearning models, including Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM) andfour machine learning models as K-Nearest Neighbors (KNN), MultilayerPerceptron (MLP), Adaptive Boosting (AdaBoost), Extreme GradientBoosting (XGB) to investigate the performance of SFP models on theBugHunter dataset. We employ the Lasso method for feature selectionand apply the Synthetic Minority Oversampling Technique (SMOTE) toaddress the issue of imbalanced data, aiming to enhance the accuracyof the results. The experimental findings reveal that CNN and RNNoutperformed other machine learning models, achieving the best overallperformance.