Software Defect Prediction Based on Double Traversal AST
Shaoming Qiu, E Bicong, Xinchen Huang, Liangyu Liu · 2024
Software defect prediction uses defect-related features to predict possible future defects in software. Abstract syntax tree (AST) is a deep semantic feature of software, which represents the structure and semantic information of software. Therefore, many studies use AST for software defect prediction, but due to the complex structure of AST, it is easy to cause information loss when extracting features, resulting in low prediction performance. Therefore, we propose a software defect prediction method DTA (Double Traversal AST and self-attention) using a double traversal AST representation method and a self-attention mechanism. This method uses two different traversal methods: root first and leaf first, to traverse the AST respectively, and input the progressive structure and parallel structure of the AST into the global word embedding for learning. After word embedding, the weights of different features are determined by the self-attention mechanism, and finally input into the convolutional neural network (CNN) to learn and predict defect tendencies. In order to compare the performance of the model, we select 21 Java open source projects and conducted cross-version and cross-project experimental comparisons with multiple models based on CNN, long short-term neural network (LSTM) and graph neural network (GNN). Experiments show that DTA improves F1 and AUC performance by 8%-14% across versions and 12%-14% across projects compared to other software defect prediction models.