A Benchmark Dataset with Larger Context for Non-Factoid Question-Answering over Islamic Text
Faiza Qamar, Seemab Latif, Nor Shahida Mohd Jamail, Rabia Latif · Data Intelligence · 2025
Accessing and comprehending religious texts, particularly the Quran (the sacred scripture ofIslam) and Ahadith (the corpus of the sayings or traditions of the Prophet Muhammad), in today’sdigital era necessitates efficient and accurate Question-Answering (QA) systems. Yet, thescarcity of QA systems tailored specifically to the detailed nature of inquiries about the QuranicTafsir (explanation, interpretation, context of Quran for clarity) and Ahadith poses significantchallenges. To address this gap, we introduce a comprehensive dataset meticulously crafted forQA purposes within the domain of Quranic Tafsir and Ahadith. This dataset comprises a robustcollection of over 73,000 question-answer pairs, standing as the largest reported dataset in thisspecialized domain. Importantly, both questions and answers within the dataset are meticulouslyenriched with contextual information, serving as invaluable resources for training and evaluatingtailored QA systems. However, while this paper highlights the dataset’s contributions andestablishes a benchmark for evaluating QA performance in the Quran and Ahadith domains, our subsequent human evaluation uncovered critical insights regarding the limitations of existing automaticevaluation techniques. The discrepancy between automatic evaluation metrics, such asROUGE scores, and human assessments became apparent. The human evaluation indicated significantdisparities: the model’s verdict consistency with expert scholars ranged between 11% to20%, while its contextual understanding spanned a broader spectrum of 50% to 90%. These findingsunderscore the necessity for evaluation techniques that capture the nuances and complexitiesinherent in understanding religious texts, surpassing the limitations of traditional automatic metrics.