Software Bug Prediction Using Machine Learning Algorithms: An Empirical Study on Code Quality and Reliability

Elavarasi Kesavan · International Journal of Innovations in Science Engineering and Management. · 2025

This study examines the effectiveness of a hybrid Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) model for predicting software bugs, with the objective of improving code quality and dependability. The research leverages the JM1 dataset from the PROMISE Software Engineering Repository and utilizes sophisticated preprocessing approaches, including Borderline-SMOTE, SMOTETomek, RobustScaler, Yeo-Johnson transformation, and Recursive Feature Elimination, to mitigate class imbalance and feature redundancy. The CNN-LSTM model attained a validation accuracy of 98.10%, a precision of 91.42%, and a recall of 99.53%, exhibiting a minimal false negative rate and indicating great sensitivity in detecting defect-prone modules. The results underscore the model's capacity to identify spatial and sequential patterns in software metrics, providing a reliable instrument for early fault identification. This work enhances software engineering by confirming deep learning's efficacy in defect prediction, offering practical insights for developers, and delineating future research avenues for cross-project generalization and model optimization.

Read the paper · More papers on PaperTik