Content based Detection and Blocking of Spam/Phishing Emails using Machine Learning

Akalya Devi C, Karthika Renuka D, S. Sarvesh · Zenodo (CERN European Organization for Nuclear Research) · 2022

Utilising the web has been increasing day by day, as a greater number of people are using it, especially for communication. E-mail remains to be one of the most efficient ways of communication techniques and one of the most effective tools for communication for social to business purposes, due to its cost and minimum time consumption. Through e-mail, one can flood the internet by sending multiple copies of same message to large number of users. One important issue to be addressed in e-mails is that our inboxes are generally affected by attacks which mainly includes spam. Currently, spam e-mails are identified by detecting stop words in it, however if any new spam, fake or irrelevant e-mail is sent without including the stop words, it isn't properly identified. Therefore, a system should learn the words and its meaning to detect spam e-mails efficiently. To overcome this issue of blocking new and unrecognised spam e-mails, Machine Learning based approach on ‘Phishing Websites’ dataset from the UCI repository is proposed. Our proposed methodology is to use Morphological Analysis in Natural Language Processing (NLP) for better spam identification. By utilising the machine learning techniques efficiently, spam and phishing emails are to be detected and blocked in the server side itself.

Read the paper · More papers on PaperTik