Effects of applying stemmers with word representation on Arabic text classification performance using deep learning methods
Mohammed Bahbib, Majid Ben Yakhlef · 2024
Due to the increasing of using the Arabic language on the internet in the last decade, it became necessary to create an automated Arabic text classification model to manage this huge number of generated texts. Many studies have investigated the ap-plications of Natural Language pre-processing (NLP) techniques in Arabic text processing. This includes stemming, word represen-tation, and the creation of classifier models. Where all contribute to improving Arabic language pre-processing and understanding. In this work, we studied the effect of different stemming methods with word representations on Arabic text classification using RNNs models. We made this study by comparing the accuracy of different combinations of stemmers aproches (without, root-based, and light-based stemming), word representation (Word2Vec, FastText, and GloVe), and four Recurrent Neural Networks (RNNs) architectures which are: LSTM, Bi-LSTM, GRU, and Bi-GRU. The best results in the terms of accuracy was achived by using the combination of no stemming with Word2Vec as a word representation and Bi-GRU as a classifier.