Text Classification of News Articles Using Machine Learning on Low-resourced Language: Tigrigna
Awet Fesseha, Shengwu Xiong, Eshete Derb Emiru, Abdelghani Dahou · 2020
Text categorization or Textual document is a method that becomes more significant in tagging a textual document to their most relevant label. However, not all languages have parallel textual growth; without free and absences of a dataset, text categorization becomes interesting for Tigrigna language, i.e., low-resourced language. Our aim to identify the given document to its categories based on its linguistic features. To achieve our goal, we have constructed a new dataset from different Tigrigna news sources. The dataset has six main categories: Agriculture, Sports, Health, Education, Religion, and Politics. Each collected is article preprocessed from Latin characters, punctuations, and stop words. We deployed a collection of different classical machine learning classifiers to investigate its effectiveness in our datasets. Namely, 7 popular classifiers were used, Logistic Regression, Nearest Centroid, Decision Tree (DT), Support Vector Machines (SVM), K-nearest neighbors (KNN), Random Forest Classifier, and Multi-Layer Perceptron (MLP). Ensemble models also implemented to get the best accuracy by combining the best classifiers based on their majority-voting classifiers. Our experimental results showed reliable performance with a minimum F1-score of 89.1% achieved by Nearest centroid and top performance of 96 % achieved by SVM. The experimental results presented in terms of precision, Recall, and F1-scores.