Investigating the Effect of Diacritics on Arabic Speech Models

Mohammed, Ibrahim Ali Ibrahim Ali · MBZUAI iRep · 2025

Arabic is a morphologically rich language in which short vowels, indicated by diacritics, play a crucial role in determining both pronunciation and meaning. However, these diacritics are frequently omitted in written Arabic, as native speakers can infer them from context. This omission introduces significant ambiguity for both human readers and NLP systems, particularly in tasks such as Automatic Speech Recognition (ASR) and Text-to-Speech synthesis (TTS), where accurate pronunciation is essential. Modern Arabic NLP models are typically trained on undiacritized text, limiting their ability to model the full phonological structure of the language. In this thesis, we address the problem of diacritic omission by explicitly integrating diacritics into the pretraining of multimodal transformer models. We investigate whether diacriticaware pretraining can improve performance, especially for Arabic speech synthesis. Our contributions include a detailed study of ASR-based diacritization, an evaluation of data augmentation strategies, and the development of a Diacritized Arabic Text and Speech Transformer (ArTST). Comprehensive evaluations across diverse Arabic domains demonstrate that the diacritic-aware ArTST model consistently achieves the lowest ASR error rates among state-of-the-art systems. Introducing random diacritic augmentation has a negligible effect on performance, neither significantly improving nor degrading results. In TTS preference tests, including diacritics during fine-tuning delivers substantial gains in naturalness and intelligibility, while an additional diacritic-aware pretraining phase yields a modest but consistent further improvement over fine-tuning alone.

Read the paper · More papers on PaperTik