Building a Twitter Social Media Network Corpus for Libyan Dialect
Husien Albashir Alhammi, Ramadan Alsayed Alfard · International Journal of Computer and Electrical Engineering · 2018
We present a new Libyan dialect corpus which contains data set of 5000 statements that written by Libyan Twitter's users as tweets.The corpus is manually classified into fifteen categories.However, the statistical results of corpus data show that there is the significant variety of tweets numbers among categories.Obviously, Libyan Twitter's users tend to express most their tweets in sports whereas the lowest number of tweets was in films category.Yet no corpus of Libyan dialect has emerged.Therefore, building such a corpus is very crucial to be used for the public purposes of not only linguistic but also NLP research for Libyan dialect.