Cue-based bootstrapping of Arabic semantic features

Khaled Elghamry, Rania Al-Sabbagh, Nagwa El-Zeiny · 2008

Motivated by the fact that semantic features are understudied in Arabic Natural Language Processing (ANLP) in spite of being essential for some Natural Language Processing (NLP) tasks such as Anaphora Resolution (AR), Word Sense Disambiguation (WSD) and Prepositional Phrase (PP) attachment, this paper presents a cue-based algorithm to build an Arabic lexicon that tackles such semantic features. The lexicon, whose entries are extracted from the World Wide Web (WWW) using bilingual and monolingual cues, achieves a performance rate of 89.7% measured according to a gold standard set of 3000 entries. Moreover, using such a lexicon raises the performance of an AR algorithm for Arabic generic corpora from 74.4% to 87.4% which is a state-of-the-art performance rate. To the best of the authors’ knowledge, this paper presents the first attempt to deal with Arabic semantic features beyond the features of gender and number.

Read the paper · More papers on PaperTik