Novel Approach towards Arabic Question Similarity Detection
Mohammad Daoud · 2019
In this paper we are addressing the automatic detection of Arabic question similarity, which is an essential issue in a variety of NLP/NLU applications such as question answering systems, virtual assistants, chatbots ... etc. We are proposing and experimenting a rule-based approach that relies on lexical and semantic similarity between questions with the utilization of supervised learning algorithms. Our approach categorizes questions semantically according to their type and scope; this categorization is based on hypothetical rules that have been validated empirically, for example, a Timex Factoid question (a question asking about time) is less likely similar to an Enamex Factoid question (a question asking about a named entity). This article details the procedures of question pairs preprocessing, lexical analysis, feature extraction and selection and most importantly the similarity measures. According to the experiment we have conducted, our approach achieved promising precision and accuracy based on a test data of 1450 question pairs.