Fault Prediction Unveiled: Analyzing the Effectiveness of RandomForest, LogisticRegression, and KNeighbors
Sai Krishna Gunda · 2024
This research explores the application of three distinct machine learning algorithms—Random Forest, Logistic Regression, and KNeighbors—to predict software faults. These algorithms are evaluated using various performance metrics, including Accuracy, F1 Score, Precision, Recall, and ROC AUC, to assess their effectiveness in identifying software defects. Random Forest, known for its ensemble learning approach, combines multiple decision trees to improve prediction accuracy. Logistic Regression, a statistical model, is used for binary classification tasks, offering insights into the likelihood of faults occurring. KNeighbors, a non-parametric method, classifies instances based on their proximity to other data points. Each algorithm is scrutinized for its ability to accurately forecast software issues, providing a comprehensive evaluation of how well it can differentiate between faulty and non-faulty instances. The comparative study reveals the strengths and limitations of each model. Random Forest may excel in handling complex relationships and interactions between features, while Logistic Regression provides a straightforward, interpretable approach. KNeighbors can be effective in capturing local patterns but may struggle with large datasets. Understanding these nuances helps in selecting the most appropriate model based on specific benchmark evaluation cases, which is critical for enhancing fault prediction performance. This research underscores the importance of choosing the right model for software defect prediction. The findings indicate that different algorithms offer varying advantages depending on the context and evaluation criteria. Consequently, future work in this area should consider these insights to refine fault prediction methods further. The paper provides a foundational guide for future research, highlighting the need for tailored model selection and evaluation strategies to improve software defect prediction accuracy and effectiveness.