Modeling Spammer Behavior: Artificial Neural Network vs. Naive Bayesian Classifier
Saiful Islam, Rafiqul Islam · InTech eBooks · 2011
Application 242 recipient will miss legitimate e-mail from a known or expected correspondent with a heretofore unknown address, such as correspondence from a long-lost friend, or a purchase confirmation pertaining to a transaction with an online retailer.A detail explanation of these techniques is given in (Islam & Chowdhury, 2005).Ramachandran et al. (2007) propose a new technique called behavioral blacklisting, which complements existing blacklists by categorizing spammers based on how they send email, rather than the IP address (or address range) from which they are sending it.The intuition behind their idea is that, while IP addresses are ephemeral as identifiers, spam campaigns, spam lists, and spamming techniques are more persistent.If one can identify email-sending patterns that are characteristic of spamming behavior, then she can continue to classify IP addresses as spammers even as spammers change their IP addresses.Machine learning algorithms namely Naïve Bayesian classifier, Decision Tree induction, Artificial Neural Network and Support Vector Machines, based on keywords or tokens extracted from the e-mail's Subject, Content-Type Header and Message Body, have been used successfully in the past (Aery & Chakravarthy, 2005 ;Drucker et al., 1999;Eichler, 2005;Islam & Chowdhury, 2005).Very soon they fall short to filter out spam emails as the spammer changing themselves in the ways that are very difficult to model by simple keywords or tokens (Stuart et al., 2004).The tactics the spammer uses follow patterns and these behavioral patterns can be modeled to combat spam.Actually the more they try to hide, the easier it is to see them (Stuart et al., 2004).Now the question is: Ques 1. Are the patterns that the spammers follow common to all?Ques 2. If the spammers follow patterns to spread spams, is it possible to track those patterns?Ques 3. If one can track the common spammer patterns, is it possible to model them?Ques 4. Is it possible to model common spammer patterns by machine learning approaches?Ques 5. What level of accuracy is possible to achieve if one apply machine learning approaches?Many researchers observe that spammers follow patterns.These patterns can be discovered from many different places: from email corpus (Stuart et al., 2004) by analysing their contents, network-level behavioral patterns (Ramachandran et al., 2006;Sperotto et al., 2009), transfer pattern during transfer sessions (Zhang et al., 2006), resource usage patterns (Xu et al., 2010) and spammer behavior in terms of the chain of machines they use to deliver their messages (Guerra et al., 2009).This study investigates the possibilities of modeling spammer behavioral patterns instead of vocabulary as features for spam email categorization and these behavioral patterns are discovered by analysing email corpus Subject, Content-Type Header and Message Body.The two machine learning algorithms Naïve Bayesian Classifier and Artificial Neural Networks are experimented to model common spammer patterns and both of them achieve a promising detection rate that can be considered as an improvement of performance compared to the keyword-based contemporary filtering approaches.