IFF-WAV2VEC: Noise Robust Low-Resource Speech Recognition Based on Self-supervised Learning and Interactive Feature Fusion
Jing Cao, Zhaopeng Qian, Chongchong Yu, Tao Xie · 2023
In recent years, self-supervised learning representation (SSLR) has shown remarkable performance in low-resource speech recognition. However, it lacks consideration for the robustness of low-resource models in noisy environments, making it crucial to enhance their noise robustness. Speech enhancement is a commonly used denoising method, but it suffers from information over-suppression during training, leading to reduced accuracy in automatic speech recognition (ASR). To address this issue, this paper proposes an innovative Iff-wav2vec network architecture. Firstly, the network architecture integrates voice enhancement, SSLR, and ASR into one network. Secondly, this article uses interactive feature fusion methods to fuse noise features and enhanced features to compensate for the lack of information in the enhanced features. Finally, experimental results on Tujia and Shui languages show that the proposed method can effectively improve low resource ASR performance under various noise settings, resulting in stronger noise robustness.