Comparative study of classifiers to find effect of structured and unstructured field of an email on accuracy of classifiers

Priti Kulkarni · 2015

E-mail has become one of the most ubiquitous methods of communication. It is used not only for personal contacts but also for business, advertising, and electronic commerce. The volume of email messages received per day varies from the tens for a regular user to the thousands for enterprises. The increasing volumes of unwanted emails, increasing the size of inbox are basic hindrances in managing emails. Therefore effective information management is a key for achieving business success. Email classification is applied on receiving emails so that emails can be properly filtered.When classifying email, data contain in messages are very complex, multidimensional, and represented by a large number of features. Directly applying classification makes learning intractable. The objective of this study is to find the effect of header (structure part) and body (unstructured part) of an email fields on the accuracy of classifier in a controlled setting. Keeping educational organizations in mind email classification strategies are studied and reported in this paper. We have compared four classification techniques K-nearest neighbour, Decision tree and Naive Bayes, Support vector machine. These algorithms are selected because of its effectiveness in text classification. Our experiment shows that both header (structure part) and body (unstructured part) of email contain important information and plays vital role in determining the class of email.

Read the paper · More papers on PaperTik