The Impact of Data Preprocessing on the Performance of a Naive Bayes Classifier

Priyanga Chandrasekar, Kai Guo Qian · 2016

In the research of text mining, document classification is a growing field. Even though we have many existing classifying approaches, Naïve Bayes Classifier is simple and effective at classification. Data preprocessing is the important step in the data mining process. It prepares the raw data for the further process. The aim of this paper is to identify the impact of preprocessing the dataset on the performance of a Naïve Bayes Classifier. The Naïve Bayes Classifier is suggested as the most effective method to identify the spam emails. The Impact of preprocessing phase on the performance of the Naïve Bayes classifier is analyzed by comparing the output of both the preprocessed dataset result and non-preprocessed dataset result. The test results show that combining Naïve Bayes classification with the proper data preprocessing can improve the prediction accuracy.

Read the paper · More papers on PaperTik