Arabic News Classification and Generation Based on an Encoder-Decoder Transformer Model (ArabicT5)

Hezam Gawbah, Nashwan Ahmed Al-Majmar, Akram Alsubari, Muneer Hazaa Alsurori, Rashed Ali Ahmed Al-Shaebi · 2024

The proliferation of Arabic information on the internet has led to a surge in the demand for automated Arabic text classification and generation techniques. Arabic text classification and generation techniques are fundamental tasks in natural language processing (NLP), which involves classifying Arabic text into pre-defined classes and Arabic text generation. However, few studies focus on Arabic text classification and generation due to the challenges in processing and understanding the Arabic language. Recent developments in the natural language processing for Arabic, particularly the emergence of transformers, paved the way for computers to comprehend Arabic more effectively and solve complex related tasks. In this paper, we tried to collect a dataset from the Youm7 news website, which we named NADCG as an abbreviation for the New Arabic Dataset for Text Classification and Generation. NADCG contains a vast number of Arabic news articles within eight categories (politics, economics, sports, health, technology, culture, arts, and accidents). We fine-tuned an Arabic-pre-trained transformer-based model called ArabicT5 on a large Arabic news corpus. Before fine-tuning, we reprocessed and cleaned data, but because NADCG size was big and required very high memory and processing, we used a sub-dataset for training. On the SANAD corpus, we fine-tuned the ArabicT5 model for classification, and on the NADCG corpus, we fine-tuned the model for classification and generation. In this experiment, the results showed that the proposed model based on an Encoder-Decoder transformer model (ArabicT5) on the NADCG corpus achieved an accuracy of 96.17% for classification and an accuracy of 87.16% for generation. Also, the proposed model based on an Encoder-Decoder transformer model (ArabicT5) improved accuracy for classification and achieved more accuracy of 96.49% for classification on the SANAD corpus, while HANGRU and CGRU, achieved accuracies of 95.81 and 93.43 consecutively on the SANAD corpus.

Read the paper · More papers on PaperTik