WiHArD: Wikipedia Based Hierarchical Arabic Dataset for Text Classification

Djelloul Bouchiha, Abdelghani Bouziane, Noureddine Doumi, Farouk Omar Berbouchi, Aymen Abdelghani Kebir, Nihad Mebarki, Badiâ Achouak Benameur · 2024

Text classification assigns a text to its corresponding category (class). It is a Natural Language Processing (NLP) task that can be performed using Artificial Intelligence (AI) methods, notably Machine Learning (ML) and deep learning. To be trained, supervised machine learning algorithms need a labelled dataset (corpus) as input. In this paper, we create a labelled dataset called WiHArD (Wikipedia based Hierarchical Arabic Dataset) and make it publicly available for the benefit of NLP and AI communities, especially those working on the Arabic language. Unlike existing datasets, WiHArD has a hierarchical structure based on the Wikipedia hierarchy. Experiments have been carried out to show the consistency and efficiency of the WiHArD dataset when used for the Arabic text classification process.

Read the paper · More papers on PaperTik