Using Twitter to Collect a Multi-Dialectal Corpus of Arabic
Hamdy Mubarak, Kareem Mohamed Darwish · 2014
This paper describes the collection and classification of a multi-dialectal corpus of Arabic based on the geographical information of tweets.We mapped information of user locations to one of the Arab countries, and extracted tweets that have dialectal word(s).Manual evaluation of the extracted corpus shows that the accuracy of assignment of tweets to some countries (like Saudi Arabia and Egypt) is above 93% while the accuracy for other countries, such Algeria and Syria is below 70%.