A Feature Subset Selection In E-mail Spam Detection
Sivannarayana Nerella, K. Rajani Devi · 2012
In the recent in formation industry huge amounts of data is being collected and store continuously. For frequent data updating data collecting agents may communicating with various sources from different places. Likely electronic mail communication is indispensable nowadays, but the electronic mail spam problem continues growing drastically. In present trends ’ the notion of collaborative spam filtering with near-duplicate similarity matching scheme has been widely discussed. The basic idea of the similarity matching scheme for spam detection is to maintain a known spam database, information passed by the user, to block subsequent near-duplicate spams. on purpose of achieving effective similarity matching and reducing storage utilization, prior works mainly represent each electronic mail by a succinct abstraction derived from mail content text. However, these abstractions of mails cannot fully catch the evolving nature of spams, and are thus not effective enough in near-duplicate detection. in this paper, we propose a new electronic mail abstraction scheme, which considers e-mail layout structure to represent electronic mails. we represent a procedure to generate the electronic mail abstraction using html content in electronic mail, and this newly devised abstraction can more effectively capture the near-duplicate phenomenon of spams. Moreover, we design a complete spam detection system cosdes (standing for collaborative spam detection system), which possesses an efficient near-duplicate matching scheme and a progressive update scheme. the progressive update scheme enables system cosdes to keep in most up-to-date information for near-duplicate finding. we evaluate cosdes on a live data set gathered from a real electronic mail server and show that our system outperforms the prior approaches in finding results and is applicable to the present world. Index terms Spam detection, electronic –mail abstraction, near-duplicate matching.