Exploring the Effects of Dimensionality Reduction Techniques on Topic Classification of an Arabic News Article Dataset

Mouaad Errami, Mohamed Amine Ouassil, Rabia Rachidi, Mohammed Jebbari, Bouchaib Cherradi, Abdelhadi Raihani · 2025

this study examines the difficulties of classifying text in Arabic using advanced machine learning (ML) algorithms with dimension reduction methods like Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP). Keeping in view the unique complexities of text in news articles written in Arabic, we employed diverse ML algorithms like Support Vector Machines (SVM), Logistic Regression (LR), Multinomial Naïve Bayes (MNB), and Random Forest. In comparative research, we examine the effect of PCA and UMAP on model performance with regard to accuracy and processing time. The findings indicate that PCA increases accuracy in all models with a maximum accuracy of 87.23% using SVM with PCA. Along with this, PCA reduces processing time significantly compared to processing raw text and is thus a good candidate to consider in text classification in large datasets. This research not only emphasizes the importance of dimension reduction in text classification in Arabic but also offers insights to enhance ML workflows in other languages with complex structures.

Read the paper · More papers on PaperTik