Learning Adapters for Code-Switching Speech Recognition
Chun-Yi He, Jen‐Tzung Chien · 2023
Multilingual code-switching speech recognition has been an emerging research direction in real-world applications since most of speakers are bilingual or multilingual. A code-switching sentence is the mixing of two or more languages especially within the same sentence. It is crucial to work out a multilingual speech recognition via a code-switching scheme based on a parameter-efficient learning by utilizing a pretrained encoder. Using this scheme, it is essential to identify the languages within a single spoken utterance for code-switching speech recognition. However, collecting speech data from monolingual language is easier than collecting code-switching speech in multiple languages. This study develops a new code-switching Mandarin-English speech recognition by utilizing a large-scale pre-trained backbone model containing 53 single-code languages. The fine-tuning of backbone model is implemented by introducing the controllable language or task adapters and incorporating small number of Mandarin-English code-switching speech where the backbone model is frozen and only individual adapters for Mandarin and English are estimated. A limited amount of controllable parameters can be sufficiently calculated. Experiments on code-switching speech recognition for Taiwanese Mandarin and English shows the merit of the proposed method.