Multi-label code smell Detection Method based on CodeBERT
Hongming Bao, Le Wei · 2024
Code smells usually hint at possible problems or design flaws in the code. Efficient code smell detection methods can significantly improve code quality. However, existing code smell detection methods rely heavily on information such as metrics and indicators, fail to make full use of the syntax and semantics of source code, and have shortcomings in multi-label code smell detection. To this end, this paper proposes a Multi-Label Code Smell Detection Based on CodeBERT (MLCSDCB) based on the pre-trained model CodeBERT. Firstly, the constructed data set is preprocessed and standardized to reduce noise interference, and the multi-label data set is constructed by the label power set method. Then, CodeBERT is used to embed the code and extract its rich syntactic and semantic information. Then, Bi-LSTM was used to fully learn the Code Representation Information [CLS] vector of each layer of CodeBERT to enhance the context understanding ability of the model. Finally, the attention mechanism was used to strengthen the key information according to the importance of each layer [CLS] vector, and the multi-label prediction was performed based on the strengthened feature vector to improve the detection effect of the model. Compared with the existing detection methods based on source code representation, MLCSDCB improves the precision, recall and F1 value by 2.02%, 1.04%and 4.75%respectively.