A framework for Sudanese Arabic – English Mixed Speech Processing
Mohammed O. Elfahal, Mohammed Elhafiz Mustafa, Mohammed Elhafiz Mustafa, Rashid Abdelhaleem Saeed · 2020 International Conference on Computing and Information Technology (ICCIT-1441) · 2020
Using more than language in a single discourse is a practical phenomenon appears in bilingual and multilingual communities known as mixed speech communication. The mixed speech occurs in communication to express ideas and thoughts using vocabulary of all used languages. This paper addresses the problem of automatic speech recognition and language identification for Sudanese Arabic and English languages in mixed speech mode, in which number and location of languages are not previously known. Native language in mixed speech is dominant based on assumption that native speaker does not suddenly reconfigure his articulation organs to produce sounds when switching to other language. This study proposes a generalized framework of Automatic Speech Recognin in Mixed Speech mode (ASR-MS). Proposed framework defines mixed speech as a hybrid language not belong to any language participates in mixed sentence. This new language is processed based on its distinct attributes not based on attributes of languages originated from. A single Language Model (AM) is built based on Sudanese Arabic pronunciation. The supporting components such as phonetic dictionary and language lexicon are bilingual, keeps words in their original forms. For measuring the soundness of a framework, a 100 Sudanese Arabic - English daily life mixed sentences were collected and recorded. Preliminary results are encouraging towards building a generalized mixed speech recognition and language identification based on speakers instead of languages.