Enhancing Spam Detection with GANs and BERT Embeddings: A Novel Approach to Imbalanced Datasets

Adnane Filali, El Arbi Abdellaoui Alaoui, Mostafa Merras · Procedia Computer Science · 2024

In recent years, the prevalence of imbalanced datasets has posed significant challenges to traditional machine learning models. This imbalance is especially pronounced in fields such as spam detection, where malicious or unwanted messages are typically outnumbered by legitimate ones. Although various techniques have been developed to address this disparity, most conventional methods either undersample the majority class or oversample the minority class, potentially leading to information loss or overftting. In this study, we propose a novel approach using Generative Adversarial Networks (GANs) to generate synthetic samples, thus enhancing the representation of the minority class. By leveraging the powerful BERT embeddings to capture the intricate textual nuances, our model strives to produce synthetic spam messages that are not only realistic but also diverse. Initial results indicate that our GAN-augmented model offers a noticeable improvement in detecting spam messages compared to traditional techniques. This advancement not only holds potential for spam detection but also suggests broader applicability in addressing dataset imbalance across various domains.

Read the paper · More papers on PaperTik