Deep Learning-Based Software Defect Prediction via Semantic Key Features of Source Code, Handling Imbalanced Datasets

Hiro Gaspar Inglês de Andrade · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT) · 2025

This work is part of the master’s thesis in Computer Engineering at the University of Beira Interior. It addresses themes related to software defect prediction, known as SDP, with the main objective of developing a predictive model using contextual features generated through deep learning models. To achieve the defined goals, five fundamental steps were followed: data preprocessing, mapping and embedding of tokens, extraction of contextual information, handling of datasets with class imbalance, and building the machine learning model for defect prediction. The dataset used was PROMISE, which encompasses software projects developed in Java, with multiple versions for each one. The experiments were conducted individually for each version, using static and contextual features generated through LSTM networks. The models were evaluated based on AUC, Accuracy, MCC, Recall, and Precision metrics. In general, it was observed that the use of contextual features resulted in significantly better performance. Among the models tested, Logistic Regression proved to be the most effective, demonstrating the best predictive capability. However, when combining different versions of the projects, a drop in performance was recorded, with the MCC showing low values, especially in the case of Naive Bayes, which in some scenarios even presented negative values. This phenomenon can be explained by factors such as concept drift (the change in data behavior over time) and overfitting (when the model fits excessively to the training data, compromising its ability to generalize), issues that have not been deeply addressed but are considered for future work.

Read the paper · More papers on PaperTik