MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)
Dhaou Ghoul, Gaël Lejeune · 2019
We present MICHAEL, a lightweight method developed for the MADAR shared task on travel domain Dialect Identification (DID).It uses character-level features and perform classification without any pre-processing.Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier.MICHAEL achieved an official score (accuracy) of 53.25% with 1 ≤ N ≤ 3 but showed a much better result with character 4-grams (62.17%).