A Rich and Balanced MSA Corpus for HMM-Based Consonant-Vowel Segmentation

Youssef Boutazart, Abderrahim Ezzine, Hassan Satori · 2024

This research delves into the creation of an innovative Modern Standard Arabic corpus, aiming for a comprehensive balance and richness while adhering to Zipf’s law. Building a phonetically diverse Arabic sentence collection yields significant advantages in terms of efficiency, cost-effectiveness, and storage capacity compared to conventional corpora. The corpus currently contains 527 sentences from various resources, the word forms generated are rich and balanced, and it contains all phonemes of the Arabic language. The sentences and expressions verify the laws of balanced phonetic distribution. We integrate with this study Hidden Markov Model (HMM) applying to analyze the underlying structure of created corpus. Indeend, HMM method is used to confirm consonant - vowels segmentation, in the corpus. We train our system by re-estimation of HMM parameters to have the maximum probability of observing Arabic symbols using Forward, Backward and Baum-Welch Algorithms. The obtained results indicate that, the distinction between consonants and vowels is a statistically significant feature inherent in the Arabic language.

Read the paper · More papers on PaperTik