Misactivation-Aware Stealthy Backdoor Attacks on Neural Code Understanding Models
Xiaobing Sun, Yiran Xiao, Lili Bo, Weisong Sun, Xiangyue Liu, Bin Li, Jiale Zhang · IEEE Transactions on Software Engineering · 2025
Neural code models (NCMs) play a crucial role in helping developers solve code understanding tasks. Recent studies have exposed that NCMs are vulnerable to several security threats, among which backdoor attack is one of the toughest. It is usually achieved through data poisoning. Specifically, backdoored NCMs work normally on the clean example but produce attacker-expected output on the example injected with backdoor triggers. However, existing backdoor attacks against NCMs face two significant drawbacks: 1) lack of stealthiness, that is trigger tokens are easily detected by defense techniques/humans when they appear in excessive numbers; 2) damage to the model’s normal performance, that is partial trigger tokens may frequently appear as benign features in the clean samples, resulting in clean samples containing them may falsely activate the backdoor. To address these drawbacks, we propose a misactivation-aware stealthy backdoor attack against NCMs through data poisoning called MISNCM. MISNCM features target-biased trigger generation, thus achieving stealthy backdoor attacks. Moreover, we utilize misactivation-aware data poisoning to create calibration samples with partial trigger tokens to reduce false activations and ensure the regular performance of the model. We conduct comprehensive experiments to evaluate the effectiveness of MISNCM in attacking NCMs used for three code understanding tasks: defect detection, clone detection, and authorship attribution. The experimental results demonstrate that the triggers generated by MISNCM achieve an average attack success rate increase of 12.67% over IR and 8.38% over AFRAIDOOR. Furthermore, MISNCM achieves a 3.64% improvement in F1 score on the code clone detection task, and an average of 5.91% improvement in accuracy on the defect detection and authorship attribution tasks, compared with the two baselines.