Adapting OpenAI’s Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets

Yuhang Yang, Yizhou Peng, Hao Huang, Eng Siong Chng, Xionghu Zhong · 2024

This paper reports on SOTA results achieved using openAI’s Whisper model with adaptation on different adaptation corpus sizes for two established code-switch Mandarin/English corpus - namely SEAME and ASRU2019 corpora.Two key experiments were conducted: a) using adaptation data from 1 to 100/200 hours to demonstrate the effectiveness of adaptation, b) examining different language ID setups on Whisper prompt. The Mixed Error Rate results show that the amount of adaptation data may be as low as 1 ~ 10 hours to achieve saturation in performance gain (SEAME), while the ASRU task continued to show performance with more adaptation data (>100 hours). For the language prompt without adaptation, the results show that various prompting strategies produce different outcomes. However, after adaptation, the Whisper model uniformly improves its performance, and language prompt becomes not critical. We believe that these results can help researchers study the adaptation of Whisper to other code-switch Languages.

Read the paper · More papers on PaperTik