Spam filtering based on PV-DBOW model

Ghizlane Hnini, Anass Fahfouh, Jamal Riffi, Mohamed Adnane Mahraz, Ali Yahyaouy, Hamid Tairi · International Journal of Data Analysis Techniques and Strategies · 2021

Many feature extraction techniques have been conducted to deal with spam e-mails. However, despite their performance and efficiency, they still have a lot of weaknesses. The term frequency-inverse document frequency (TF-IDF) and the bag-of-words (BoW) are two well-known methods. Yet, they do not capture the semantic aspect of the e-mails, which may lead to misclassification. To tackle this issue, we propose an architecture based on distributed bag-of-words version of paragraph vector (PV-DBOW). It is considered as a deep learning architecture. The features generated from an e-mail are characterised by their richness, and they capture the semantic aspect of the e-mails by taking into account the context of the sentences. The obtained results show that the proposed approach outperforms the state-of-the-art methodologies in terms of precision, recall, F-measure, and accuracy.

Read the paper · More papers on PaperTik