Spam Filter by Using χ2 Statistics and Support Vector Machines
Songwook Lee · KIPS Transactions on Software and Data Engineering · 2010
ABSTRACT We propose an automatic spam filter for e-mail data using Support Vector Machines(SVM). We use a lexical form of a word and its part of speech(POS) tags as features and select features by chi square statistics. We represent each feature by TF(text frequency), TF-IDF, and binary weight for experiments. After training SVM with the selected features, SVM classifies each e-mail as spam or not. In experiment, the selected features improve the performance of our system and we acquired overall 98.9% of accuracy with TREC05-p1 spam corpus.Keywords:Spam Mail Filter, Support Vector Machine, Chi Square Statistics, Feature Selection 1. 서 론 1) 인터넷의 발달과 웹 메일 서비스의 보급으로 인해 전자우편은 그 편리함으로 인해 실생활에 널리 사용되고 있다. 그러나 인터넷의 상업적 이용과 개인정보를 이용한 범죄의 목적 등으로 매일 수신되는 스팸메일이 점점 많아지고 있다.스팸메일이란 불특정 다수에게 수신자의 동의 없이 발송되며, 수신자에게 불필요한 정보를 담고있는 전자우편을 뜻하며, 이러한 스팸메일은 사용자의 불편을 초래할 뿐만 아니라 이메일 시스템에 상당한 부하를 준다. 이러한 스팸메일을 차단하는 스팸메일 필터링에 관한 연구가 활발히 진행되고 있는데, 대부분의 연구는 베이지안 분류기를 기반으로 하고 있으며[1-5], 그 외, 마코프 랜덤 필드(Markov Random Field) 모델[6]과 k-Nearest Neighbor(k-NN) 방법[7], 최대