Building a Twitter Social Media Network Corpus for Libyan Dialect

Husien Albashir Alhammi, Ramadan Alsayed Alfard · International Journal of Computer and Electrical Engineering · 2018

We present a new Libyan dialect corpus which contains data set of 5000 statements that written by Libyan Twitter's users as tweets.The corpus is manually classified into fifteen categories.However, the statistical results of corpus data show that there is the significant variety of tweets numbers among categories.Obviously, Libyan Twitter's users tend to express most their tweets in sports whereas the lowest number of tweets was in films category.Yet no corpus of Libyan dialect has emerged.Therefore, building such a corpus is very crucial to be used for the public purposes of not only linguistic but also NLP research for Libyan dialect.

Read the paper · More papers on PaperTik