Comparative Performance of Supervised Learning Models for Software Defect Detection
Talari Nandeesh, Ashu Mehta · 2024
Predicting software defects is essential for enhancing software quality by identifying faulty components early in the development lifecycle. This research explores the effectiveness of different supervised learning algorithms in predicting defects using two well-known datasets, JM1 and KC2. The study examines the performance of five machine learning models: Logistic Regression, Gaussian Naive Bayes (GNB), K-Nearest Neighbours (KNN), Quadratic Discriminant Analysis (QDA), and Linear Discriminant Analysis (LDA). These models are assessed using important metrics like as precision, recall, F1-score, and accuracy for both defect-prone and non-defect-prone categories. Results indicate that KNN surpasses the other algorithms on both datasets, achieving the highest F1-scores for defect-prone instances, with values of 0.94 for JM1 and 0.65 for KC2. Conversely, Gaussian Naive Bayes and Quadratic Discriminant Analysis show lower performance, particularly due to reduced precision and recall in detecting defects. These outcomes underscore the difficulties caused by class imbalance in defect prediction and highlight the need for careful model selection to achieve the best results. This study provides insightful information on the strengths and weaknesses of popular classifiers used in defect prediction., providing useful recommendations for both researchers and industry professionals striving to improve defect detection systems.