Can Interpretability of Deep Learning Models Detect Textual Adversarial Distribution?
Ahoud Alhazmi, Abdulwahab Aljubairy, Wei Emma Zhang, Quan Z. Sheng, Elaf Alhazmi · ACM Transactions on Intelligent Systems and Technology · 2025
Deep Neural Networks (DNNs) are widely used in Natural Language Processing (NLP). However, adversarial samples attack benign inputs to readily fool the DNN models. The detection of these samples is a significant challenge that has received little attention in textual domains. Existing defense strategies either assume prior knowledge of specific threats or do not perform well on complex models. In this article, we provide a new framework, namely TADD for detecting textual adversarial samples by leveraging the interpretability of DNNs. In particular, we distinguish between the adversarial distribution and the benign distribution for the decision boundary of the victim models. Our method applies to NLP tasks and does not require re-training victim models and prior knowledge of adversarial attack methods. We evaluate our detector against the state-of-the-art attack methods on various real-world datasets. As demonstrated in the extensive experiments, our approach effectively discriminates between adversarial and benign samples. Additionally, our method is competitive against unseen attacks, reflecting its ability to discover new adversarial samples generated by future attack methods.