Text-Informed Knowledge Distillation for Robust Speech Enhancement and Recognition
Wei Wang, Wangyou Zhang, Shaoxiong Lin, Yanmin Qian · 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) · 2022
Most existing speech enhancement (SE) approaches heavily depend on simulated data for training, leading to performance degradation on realistic data and subsequent speech recognition task. One of the main reasons is that SE models cannot be trained on real data due to the absence of reference signals. In this paper, we aim to tackle this problem by exploiting transcribed real data to mitigate the mismatch between training and evaluation. A text-informed SE teacher is first trained to provide “reference” signals for the transcribed real data. Then a SE student is trained on both simulated and real data, where the supervision comes from the simulated ground truth and the teacher, respectively. Finally, a speech recognition model is trained on enhanced signals from the SE student. Our experimental results show that the proposed method can not only improve the speech enhancement performance, but also reduce the word error rate on the downstream speech recognition task.