Not So Naive Online Bayesian Spam Filter
Baojun Su, Congfu Xu · Innovative Applications of Artificial Intelligence · 2009
Spam filtering, as a key problem in electronic communica- tion, has drawn significant attention due to increasingly huge amounts of junk email on the Internet. Content-based fil- tering is one reliable method in combating with spammers' changing tactics. Na¨ ive Bayes (NB) is one of the earliest content-based machine learning methods both in theory and practice in combating with spammers, which is easy to imple- ment while can achieve considerable accuracy. In this paper, the traditional online Bayesian classifier are enhanced by two ways. First, from theory's point of view, we devise a self- adaptive mechanism to gradually weaken the assumption of independence required by original NB in the online training process, and as a result of that our NSNB is no longer na¨ive. Second, we propose other engineering ways to make the clas- sifier more robust and accuracy. The experiment results show that our NSNB does give state-of-the-art classification per- formance on online spam filtering on large benchmark data sets while it is extremely fast and takes up little memory in comparison with other statistical methods.