Lisan: Yemeni, Iraqi, Libyan, and Sudanese Arabic Dialect Corpora with Morphological Annotations
Mustafa Jarrar, Fadi A. Zaraket, Tymaa Hasanain Hammouda, Daanish Masood Alavi, Martin Wählisch · 2023
This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects (${\text{L}\hat{\text{i}}\text{sa}\bar{\text{n}}}$) corpora. ${\text{L}\hat{\text{i}}\text{sa}\bar{\text{n}}}$ features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The Yemeni corpus ($\tilde 1.05{\text{M}}$ tokens) was collected automatically from Twitter. The corpora of the other three dialects ($\tilde 50{\text{K}}$ tokens each) was manually collected from Facebook and YouTube posts and comments. Thirty-five (35) annotators who are native speakers of the target dialects carried out the annotations. The annotators segmented all words in the four corpora into prefixes, stems and suffixes and labeled each with different morphological features such as part of speech, lemma, and a gloss in English. We developed the Arabic Dialect Annotation Toolkit (ADAT) to assist the annotators and to ensure compatibility with SAMA and Curras tagsets. We trained annotators on a set of guidelines and on how to use ADAT. ADAT is open source, and the four corpora are available at https://sina.birzeit.edu/currasat.