Arabic Text Classification of News Articles Using Classical Supervised Classifiers

Leen Al Qadi, Hozayfa El Rifai, Safa Obaid, Ashraf Elnagar · 2019

Automatic document categorization gains more importance in view of the plethora of textual documents added constantly on the web. Text categorization or classification is the process of automatically tagging a textual document with most relevant label. Text categorization for Arabic language is interesting in the absence of large and free datasets. Our objective is to automatically identify the category of a document based on its linguistic features. To achieve this goal, we constructed a new dataset which contains almost 90k Arabic news articles with their tags from Arabic news portals. The dataset shall be made freely available to the research community on Arabic computational linguistics. The dataset has four main categories: Business, Sports, Technology and Middle East. Each collected article was cleaned from Latin characters, numbers, punctuation and stop words. To investigate the effectiveness of the dataset, we used an array of classical supervised machine learning classifiers. Namely, the following 10 popular classifiers were used: Logistic Regression, Nearest Centroid, Decision Tree (DT), Support Vector Machines (SVM), K-nearest neighbors (KNN), XGBoost Classifier, Random Forest Classifier, Multinomial Classifier, Ada-Boost Classifier, and Multi-Layer Perceptron (MLP). In pursuit of high accuracy, we implemented an ensemble model to combine best classifiers together in a majority-voting classifier. Our experimental results showed solid performance with a minimum Fl-score of 87.7%, achieved by Ada-Boost and top performance of 97.9% achieved by SVM. The experimental results are presented in terms of confusion matrices, Fl-scores, and accuracy.

Read the paper · More papers on PaperTik