Arabic Semantic Textual Similarity Identification based on Convolutional Gated Recurrent Units

Adnen Mahmoud, Mounir Zrigui · 2021 International Conference on INnovations in Intelligent SysTems and Applications (INISTA) · 2021

The augmentation of data exchanged on the internet has favored the practice of paraphrase. Its detection is one of the fundamental tasks of Natural Language Processing (NLP). It consists of identifying the degree of semantic similarity between sentences that convey the same meaning with different words. Although many researchers focused on this task on the English language, there are few works for other languages like the Arabic. It has presented important challenges because of its richness of features and processing complexities. Nowadays, deep neural networks have yielded immense success in most NLP applications. In this paper, a Siamese architecture is proposed for Arabic paraphrase detection. It has proven its relevance for semantic textual similarity. Indeed, the effectiveness of feed forward and recurrent neural networks models are studied. Experiments are conducted on the paraphrased Open-Source Arabic Corpora (OSAC). It is generated semi-automatically and validated using the benchmark SemEval. Evaluations demonstrated that Gated Recurrent Unit (GRU) outperformed Convolutional Neural Network (CNN) and other state-of-the-art methods.

Read the paper · More papers on PaperTik