A Quran and Hadith Speech Resource and Benchmark for Arabic ASR, with Professional-Reciter Training and Validation
Mohamed Kotb · Zenodo (CERN European Organization for Nuclear Research) · 2026
Automatic speech recognition (ASR) errors on sacred text are not ordinary errors: a plausible sounding substitution can alter the meaning of the Quran or of a Prophetic tradition (hadith).As chatbots, search, and summarizers increasingly answer from transcriptions rather than fromsource audio, such an error propagates as a silent, trusted corruption of scripture. Quranicrecitation is served by several speech corpora, but the hadith — the second primary sourceof Islam — has, to our knowledge, no public audio ASR resource at all; existing hadithdatasets are text-only. We introduce a unified Arabic speech resource and benchmark that (i)provides the first large hadith-audio training corpus — ~10,000 non-repeated hadith fromMajmaʿ al-Zawāʾid across 14 canonical books — together with professional-reciter Sahih alBukhari and Sahih Muslim audio as a disjoint validation benchmark, (ii) pairs both with full-Quran recitation so that both primary sources are covered under one protocol, and (iii)integrates cleaned general Modern Standard Arabic (MSA) data so that models trained on theresource remain usable for everyday transcription. Leakage control is deliberate: Quran andhadith are recited by different professionals in training and validation, and the hadith validationtext is drawn from different source books than the hadith training text, so the hadith benchmarkmeasures generalization to unseen text, unseen books, and an unseen voice. We release threedatasets — 119 h train / 44 h 47 m validation for the main set, plus a 40 h / 10 h Egyptian-dialectset and a timestamped variant — under one reproducible audio and text normalization pipeline.Baselines: a fine-tuned whisper-large-v3 reaches 0.33 % WER on Quran and 3.60 % onthe held-out Bukhari/Muslim hadith benchmark, against 3.96 %/3.99 % for CohereTranscribe Arabic, the current top open Arabic model — a ≈12× relative reduction onQuran — while retaining general-MSA performance on held-out broadcast speech. A controlled±sacred-text ablation shows the resource is additive: it can be folded into an existing trainingmix without degrading the domains a model already serves, which is what makes it practical toreduce Quran and Hadith error in systems never designed for sacred text.