Using Twitter to Collect a Multi-Dialectal Corpus of Arabic

Hamdy Mubarak, Kareem Mohamed Darwish · 2014

This paper describes the collection and classification of a multi-dialectal corpus of Arabic based on the geographical information of tweets.We mapped information of user locations to one of the Arab countries, and extracted tweets that have dialectal word(s).Manual evaluation of the extracted corpus shows that the accuracy of assignment of tweets to some countries (like Saudi Arabia and Egypt) is above 93% while the accuracy for other countries, such Algeria and Syria is below 70%.

Read the paper · More papers on PaperTik