Text Corpus and Acoustic Model Addition for Indonesian-Arabic Code-switching in Automatic Speech Recognition System
Rizky Elzandi Barik, Dessi Puji Lestari · 2019
Code-switch is a common phenomenon found during daily conversation, especially in Indonesia. In this Muslim-majority country, the occurrence of Arabic Language during conversation is quite frequent. Unfortunately, there has not been any special handling done to minimize errors caused by Indonesian-Arabic code-switching occurrences. The handling can be done on several levels. In the lexicon, Arabic vocabularies are added. Code-switching sentences are also added into the training text corpus to improve the language model. On the acoustic model, we apply Indonesian-Arabic phone merging by using the IPA (International Phonetic Association) rule to improve the performance of the acoustic model. Recognition of code-switch speech Indonesian-Arabic achieved WER (Word Error Rate) improvement of 20.87% from the baseline system. OOV rate and perplexity of the system also improved.