CodeSift: Approach for detecting source code feature redundancy

Ao Xu, Yi Zhu, Qiao Yu, Guosheng Hao · 2024

In the field of source-based software defect prediction, it is necessary to convert the source code into the data form that the model can process, which is called word embedding. The commonly used word embedding models are Word2Vec and BERT models. Meanwhile, in software defect prediction, accurate feature representation is crucial to the performance of the defect prediction model. However, feature redundancy problems in source code, such as highly similar word vectors caused by code appearing in pairs, result in a set of word vectors generated when training word embedding models that may degrade model performance when applied to defect prediction tasks. In order to alleviate this problem, a CodeSift method is proposed in this paper. By calculating the similarity of each pair of word vectors in the code word vectors generated by the word embedding model and interpolating the generated code word vectors, the problem of feature redundancy between word vectors is reduced. CodeSift generates a new word vector from a highly similar word vector, and retains the original information by assigning weights, thus improving the compactness and information richness of feature representation. Experiments show that the F1 value of the defect prediction model is improved by using CodeSift method, and the false positive rate is lower than that of the original model.

Read the paper · More papers on PaperTik