Development of anti-spam technique using modified K-Means & Naive Bayes algorithm
Devendra Kumar Tayal, Amita Jain, Kanak Meena · International Conference on Computing for Sustainable Global Development · 2016
In recent years, the issues of expanding spam mail on the web has turned into a major issue and has also become difficult to detect. Junk mails or unsolicited bulk emails are known as spam mails. They may contain malicious content and they are sent to numerous recipients through email. Spam emails also contain the malwares in executable file attachments. At commercial level, many companies hire the spammers to publicize their information regarding the offers, as it is the fastest and cheapest way of advertising. Spammers are the group of the people who apply different techniques to bypass the spam filtering methods. The general classifications of spam filtering techniques are Rule-based classification or Non Machine Learning (NML) which uses set of rules to classify whether the incoming message is spam or not. Content based classification that use machine learning techniques have given quite a promising result. Machine Learning (ML) is concerned with development of algorithms that allow computer to take intelligent decision on the basis of dataset. Some of the commonly used statistical filters are Naive Bayes, K-Means, Support Vector Machine and TF-IDF. This paper proposes a new approach to detect spam mails using linear approach of Modified K-Means & Naive Bayes classification algorithm and the Modified K-Means algorithm was proposed by Malay K. Pakhira in year 2009 to avoid empty clusters [19] which is used in our approach. This proposed approach offers the advantage w.r.t modified K-means algorithm such as improved classification accuracy, decreasing the number of iteration steps.