Efficient Learning Approach for Pronominal Anaphora and Ellipsis Identification and Resolution in Arabic Texts
Saoussen Mathlouthi Bouzid, Chiraz Ben Othmane Zribi · IEEE/ACM Transactions on Audio Speech and Language Processing · 2021
The anaphors and ellipses resolution is a challenging task for most of NLP applications. However, few works have addressed this problem in Arabic. In fact the richness of the Arabic language and the scarcity of labeled data make the task more difficult. This paper offers a generic approach that deals with the identification and the resolution of both pronominal anaphors and ellipses in Arabic texts. To identify referential anaphors and ellipses positions, we propose a self-training SVM method based on a set of pattern-based and linguistic-based criteria. For the resolution step, we propose a novel hybrid method combining a reinforcement learning method with Word Embedding models. The reinforcement learning method uses an adapted version of Q-learning algorithm to finds the optimal combination of features. It exploits a set of morphological and syntactical features. The Word Embedding based method uses word representation models to check the semantic validity of the candidates. The evaluation of the identification approach gives a precision of 99.23% for pronominal anaphors and 94.33% for ellipses. The resolution approach reaches a precision of 80% for pronominal anaphors and 97.5% for ellipses. Our proposed approaches can be easily adapted to English and French to provide multilingual approaches.