Software defect detection based on source code structuring
Hongchang Li, Wenguang Xie · 2024
Automatic detection of code defects is crucial in modern software development. Traditional methods rely on manual feature extraction and rule matching, posing challenges in addressing intricate and evolving defect types. Deep learningdriven methods for software defect detection harness neural networks to extract defect characteristics from code, thereby augmenting both the efficiency and accuracy of detection. For deep learning tasks, both data preprocessing and network model design play equally crucial roles, directly impacting the model's final prediction outcomes. Nonetheless, prevailing research predominantly focuses on refining network architectures, often neglecting advancements in data preprocessing methodologies. To yield more precise data preprocessing outcomes, we propose a multi-stage source code structuring methodology. Initially, distractions within the code text are eliminated through deduplication and normalization procedures, followed by converting the text into a standardized fixed-length vector representation using disambiguation and vectorization techniques. Drawing upon the structured representation of the source code, we employ deep neural networks to discern defects within the code. Experimental findings demonstrate the pronounced superiority of the proposed approach over traditional methods on public datasets, with notable advancements observed in precision, recall, and F1 scores. Our study highlights the vast potential of adopting structured source code representations for enhancing the efficacy of software defect detection systems.