Arabic Dialect Identification with an Unsupervised Learning (Based on a Lexicon). Application Case: ALGERIAN Dialect

Imène Guellil, Faiçal Azouaou · 2016

The identification of Arabic dialects is considered to be the first pre-processing component by any NLP problem. This identification is useful for automatic translation, information retrieval, opinion mining and sentiment analysis. Most of the work on the identification treat this issue like any classification problem. These work are based on supervised learning. The majority of them are based on the EGYPTIAN (EGY), the TUNISIAN (TUN), the IRAQI (IRAQI) dialect, etc., by omitting the Algerian dialect. The purpose of this paper is to identify the Arabic dialects within social media in an unsupervised manner. In order to do so, we use an Algerian dialect lexicon and based on Algeriers Dialect. We also propose an approach based on an algorithm that perform three kinds of identification: 1) total (when the term is totally identify), 2) partial (when the term is identified only partially with prefix and suffix) and by using the improved Levenshtein distance (based on classical Levenshtein distance) and with considering the number of character of the words compared). We applied our algorithm on a corpus of 100 messages that were collected using the Facebook API, we obtain a rate exceeding 60%.

Read the paper · More papers on PaperTik