Enhancing Performance of naïve bayes in text classification by introducing an extra weight using less number of training examples

Shahnaj Parvin Shathi, Md. Delowar Hossain, Md Nadim, Sayed Golam Rasul Riayadh, Tangina Sultana · 2016

This paper presents an effective and efficient method for classifying text documents in order to deliver feasible information retrieval using naïve bayes algorithm. Today lots of algorithms have earned good score in the field of information retrieval, Naïve Bayes is one of them. In this paper, a Weight Matrix is introduced during training text documents which is combination of term frequency (TF) and inverse class frequency(ICF) and later this weighted term is powered by a significant number and added with the posteriori value during the prediction time of Naïve Bayes (NB) algorithm to establish a better and efficient performance of the classification task. Here the precedence base element TF results an additional weight for each term (word) of the text. On the other hand, ICF gives each common word a low score. Finally the combinational term ‘Weight Matrix’ gives an extra weight and balances weight where necessary. As a result, improve the performance accuracy of the NB classifier. Experimental results show that NB with Weight Matrix rarely demotes accuracy compared to standard Naïve Bayes, instead of enhancing accuracy dramatically.

Read the paper · More papers on PaperTik