Software Bug Count Prediction Using Abstract Syntax Trees (ASTs)

Rahmeh Fawaz Ibrahim, Abdallah D. Qusef · 2024

Predicting software defects is important for ensuring high-quality software delivery.This proposed study investigates the effectiveness of deep learning models in predicting software bug counts compared to traditional machine learning models.We concentrated our investigation on two fundamental questions: (1) Can deep learning techniques surpass traditional machine learning approaches in accurately predicting bug counts in software modules via regression analysis?(2) In the context of regression-based bug count prediction, does a deep learning model that leverages features extracted from Abstract Syntax Trees (ASTs) demonstrate superior performance compared to one that relies on Object-Oriented (OO) metrics?To address RQ-1, we trained layered Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNNs) on (17) OO metrics from the PROMISE and SPSC datasets and compared their performance.For RQ-2, we extracted relevant features from ASTs using specific node types and combined them with the OO metrics to train the LSTM and CNN models.Four experiments were conducted: LSTM on OO metrics, LSTM on AST features, CNN on OO metrics, and CNN on AST features.Evaluation using Mean Absolute Error (MAE) and Mean Relative Error (MRE) showed that deep learning models outperformed traditional machine learning models, with layered LSTM and CNN models trained on combined OO metrics and AST features achieving the best defect predictive accuracy.These findings highlight the potential of deep learning approaches in software defect prediction and show the importance of feature selection in achieving optimal results.

Read the paper · More papers on PaperTik