Combining classifiers for supertagging Arabic texts
Chiraz Ben Othmane Zribi, Fériel Ben Fraj, Mohamed Ben Ahmed · 2010
This paper deals with supertagging Arabic texts with ArabTAG formalism, a semi-lexicalised grammar based on TAG and adapted for Arabic. Supertagging is a very useful task because it reduces and speeds the work of parsing. We view this problem as a classification task where elementary structures supertags (classes) are affected to words in a given sentence according to their description (morpho-syntactic and contextual information). We propose to combine three classifiers: Naïve Bayes, k-Nearest Neighbors (k-NN) and Decision tree by a voting procedure. The primary results were satisfactory as we obtained an accuracy rate of 76% although the small size of our training corpus (5,000 words) and the difficulties related to Arabic language specificities.