Probability Modeling for Improving Spam Filtering Parameters
S.C. Chiemeke, Olumide Babatope Longe · 2008
As efforts at combating electronic mail Spam rages on, users and technocrats are becoming fuzzy in their judgment of what actually constitute Spam. Most Spam filters depend on routing information, sender addresses, message title, evidence of bulk mailing, non-solicitation and mail contents to identify Spam messages. Some researchers have posited that it is the act of blindly mass-mailing a message that makes it Spam and not the actual content of the message. This to our understanding relegates the use of message content as an important factor in designing filters for identifying Spam mails. To identify the effectiveness of filtering with or without considering mail contents, we calculated the probability of the occurrence of events; Spam and real mails (called Ham) in terms of mutually inclusive and exclusive events. We then used probability theory to model parameters for identifying their instances. Our results showed that the efficiency of filtering Spam mails using other parameters and without consideration for the message content in an inclusive mail corpus reduces the volume of false positives by 8.3%. For an exclusive Spam corpus, that is, mails that are identified as Spam, filtering without consideration for message content reduced the efficiency of the filter by 7.7%.