An Analytical Study on Email Classification Using 10-Fold Cross-Validation

Takorn Prexawanprasut, Piyanuch Chaipornkaew · 2019

Start-up companies nowadays face main problems in managing a large amount of data, currently called “Big data” [1]. It is possible to say that most crucial data appear in email contents. As a result, email is considered as a valuable database. In order to obtain information from email messages, one possible preprocess is categorizing emails based on their contents. The research proposes a mechanism to classify email based on their contents. In order to evaluate the proposed mechanism, a 10-fold cross-validation method is applied. The empirical results demonstrate that the accuracy of email classification is approximately 61.89%, while the standard deviation is 2.77. The accuracy rate does not meet the research expectation perhaps because of the variance in word frequency in each category. Therefore, future work could weight each word to improve prediction performance.

Read the paper · More papers on PaperTik