Unsupervised Domain Adaptation on End-to-End Multi-Talker Overlapped Speech Recognition

Lin Zheng, Han Zhu, Sanli Tian, Qingwei Zhao, Li Ta · IEEE Signal Processing Letters · 2024

Serialized Output Training (SOT) has emerged as the mainstream approach for addressing the multi-talker overlapped speech recognition challenge due to its simplicity. However, SOT encounters cross-domain performance degradation which hinders its application. Meanwhile, traditional domain adaption methods may harm the accuracy of speaker change point prediction evaluated by UD-CER, which is an important metric in SOT. To solve these issues, we propose Pseudo-Labeling based SOT (PL-SOT) for domain adaptation by treating speaker change token ($$) specially during training to increase the accuracy of speaker change point prediction. Firstly, we improve CTC loss by proposingWeakening and Enhancing CTC(WE-CTC) loss to weaken the learning of error-prone labels surrounding$$while enhance the emission probability of$$through modifying posteriors of the pseudo-labels. Secondly, we introduceWeighted Confidence Filter(WCF) that assigns higher scores of$$to exclude low-quality pseudo-labels without hurting the$$prediction. Experimental results show that PL-SOT achieves 17.7%/12.8% average relative reduction of CER/UD-CER, with AliMeeting as source domain and AISHELL-4 along with MagicData-RAMC as target domain.

Read the paper · More papers on PaperTik