Arabic Text Categorization Using Classification Rule Mining

Mofleh Al‐diabat · 2012

Text categorization is one of the known problems in classification data mining. It aims to mapping text documents into one or more predefined class or category based on its contents of keywords. This problem has recently attracted many scholars in the data mining and machine learning communities since the numbers of online documents that hold useful information for decision makers, are numerous. However, the majority of the research works conducted on text categorization is mainly related to English corpuses and little works have focused on Arabic text collections. Thus, this paper investigates the problem of Arabic text categorization using different rule-based classification approaches in data mining. Precisely, this research works attempts to evaluate the performance of different classification approaches that produce simple “If-Then” knowledge in order to decide the most applicable one to Arabic text classification problem. The rule-based classification algorithms that the paper investigates are: One Rule, rule induction (RIPPER), decision trees (C4.5), and hybrid (PART). The results indicate that the hybrid approach of PART achieved better performance that the rest of the algorithms.

Read the paper · More papers on PaperTik