Classification of Bangla Text Documents based on Inverse Class Frequency

Ankita Dhar, Niladri Sekhar Dash, Kaushik Roy · 2018

With the increasing availability of the textual content on the internet, automatic text classification or text categorization analogously becomes a prime key in solving the problem of information organization and knowledge management. This paper aims to provide an automatic text classification system that classifies the Bangla text document into its respective classes based on the newly proposed feature selection method:TF-IDF-ICF. Here, it has been shown that the newly proposed method; TF-IDF-ICF: inverse class frequency along with TF-IDF achieves accuracy of 98.87% which is quite a satisfactory result in respect to other two feature extraction and selection methods (TF-IDF) on Bangla text documents after applying Naive Bayes Multinomial classification approach. Various other classification algorithms have also been applied for comparison purposes to test the systems performance.

Read the paper · More papers on PaperTik