Arabic speech analysis to identify factors posing pronunciation disorders and to assist learners with vocal disabilities
Naim Terbeh, Ayman Trigui, Mohsen Maraoui, Mounir Zrigui · 2016
The literature seems rich with studies addressing the detection of pronunciation disorders. The features contained in the speech signal and natural language processing techniques present famous parameters used for this objective. Despite the diversity of factors posing pronunciation disorders (vocal pathologies, non-native speakers, psychological state, age, etc.), no work has been extended to identify these factors and to assist speakers with pronunciation defects in learning spoken languages. The current work presents an original approach based on the probabilistic-phonetic modeling of Arabic speech to detect vocal disorders [1]. If the analyzed speech presents some degradations, the forced alignment score technique will be introduced to distinguish between two main factors that pose mispronunciations. Pronunciation defects can be from a native speaker suffering from vocal pathology or from a non-native speaker who learns the spoken Arabic language as an L2. Also, a platform is developed to assist speakers with degraded speeches in learning the spoken Arabic language. The present work accounts five steps. The first step consists in calculating the referenced phonetic model of the Arabic speech. This model will be used in detecting the vocal defects contained in the Arabic speech. Second, the referenced forced alignment scores for Arabic phonemes are calculated. In the third phase, for each new speaker with vocal disorders, their forced alignment scores of non-problematic phonemes are calculated [10]. In the fourth step, the two previous scores are compared to distinguish between the pronunciation disorders caused by native speakers suffering from vocal pathologies and by non-native speakers who do not master Arabic-phoneme pronunciation. The last phase consists in developing a platform to assist speakers with pronunciation defects to learn the spoken Arabic language. We are satisfied with the obtained results. We have attained an identification rate of factors posing pronunciation disorders of 95%, and the speakers using our platform have shown a good progression. Speech therapists, biologists and computer scientists can benefit from this work to develop performant systems of pathological speech processing: pathological speech recognition, accent evaluation, e-learning, etc.