Adversarial Spam Generation Using Adaptive Gradient-Based Word Embedding Perturbations
Jonathan Gregory, Qi Liao · 2023
In recent years, artificial intelligence (AI) and machine learning (ML) have become extremely promising in almost every aspect of our lives, including in cybersecurity. For instance, intrusion detection systems (IDS) and spam filters use machine learning algorithms to constantly monitor networks for abnormal behavior. However, the security of AI/ML-based solutions remains largely unknown and a cause for concerns. This study examines the possibility of cheating AI/ML-based cybersecurity solutions such as spam filters. In particular, we developed an Adaptive Gradient-based Word Embedding Perturbations (AG-WEP) framework for automatically generating adversarial spam examples. AGWEP smartly chooses the optimal perturbations across all features of the word vectors to minimize the degree of modifications to real spam messages. The experimental study suggests the adversarial model is effective to generate meaningful adversarial examples to fool a CNN-based spam classifier.