Crafting Adversarial Email Content against Machine Learning Based Spam Email Detection
Chenran Wang, Danyi Zhang, Suye Huang, Xiang‐Yang Li, Lei Ding · 2021
While machine learning based spam detectors have proven useful, spammers are learning to bypass the detectors by modifying their email content. Adversarial attacks on machine learning models have been observed in domains such as image classification. Applying such adversarial attack algorithms to craft spam emails to evade spam email detectors, however, has limitations. Such algorithms generate adversarial perturbations in the feature space. Different from image data, translating the adversarial perturbations from the feature space to text formats, as in emails, changes the effectiveness of the adversarial perturbations. It can reduce the attack success rate in the case of spam email detection. In this paper, we study the feasibility of adversarial attacks on machine learning based spam detectors and propose two novel text crafting methods leveraging adversarial perturbations generated by the adversarial example generation algorithms to improve the attack effectiveness. One method tries to approximate the feature values and the other adds special words to original emails. In experimentation, we use PGD as an example to demonstrate and compare the effectiveness of our attack methods on spam email detectors. We also examine the transferability of the proposed attack methods on different machine learning models.