Robust Text Classifier for Classification of Spam E-Mail Documents with Feature Selection Technique

Akhilesh Kumar Shrivas, Amit Kumar Dewangan, Samrendra Mohan Ghosh · Ingénierie des systèmes d information · 2021

E-mails are an effective medium for sending information in various modes like text, audio, video, etc. from one person to another.Spam e-mail is a junk e-mail that unnecessary wastage memory space, wasting time to delete and maintain e-mails in the mailbox.The contribution of this research work is to develop a robust and computational efficient classifier that classifies the spam e-mail and ham e-mail documents.This paper analyzes and validates the spam e-mails documents using different data mining-based classification techniques.The most importance of this research work is to select the best classifier with reduce feature subset of datasets that achieve better accuracy compared to other existing classifiers.We have collected six types of Enron datasets and prepared the last seven Enron datasets that combine all these six Enron datasets.Then, filtering the datasets with the help of the WEKA data mining tool.In the first step, we perform preprocessing the datasets and remove all the irrelevant words from the datasets.We have used different classifiers like Nave Bayes, J48, Random Forest, Random Tree, and Adaboosting to analyze and classify ham and spam e-mails documents.We also compare the performance of the classifier in terms of accuracy where Random Forest gives better accuracy with all seven Enron datasets.Finally, we have used the SymmetricalUncert feature selection technique to make the optimized dataset with a reduced feature subset.The suggested Random Forest classifier gives 98.73% of accuracy with reduced features of Enron datasets.

Read the paper · More papers on PaperTik