Dual-Strategy Fusion Method in Noise-Robust Speech Recognition

Jiahao Li, Cunhang Fan, Enrui Liu, Jian Zhou, Zhao Lv · 2024

Automatic Speech Recognition (ASR) systems have become in-tegral to various aspects of people's lives. However, the presence of noise in real-world environments often affects ASR performance. The mainstream approach to mitigate this issue is to use Speech Enhancement (SE) as a preprocessing step. Nevertheless, this process inevitably introduces speech distortion, which further degrades ASR performance. To address this problem, this paper proposes an end-to-end robust ASR framework with a dual-strategy fusion method for joint training. The dual-strategy fusion method consist of Mask and Mapping Fusion (MMF) and Interactive Feature Fusion (IFF). These strategies integrate different features at various stages to reduce the impact of speech distortion. We conducted experiments on the open-source Mandarin speech corpus AISHELL-1 and two noise datasets, 100 Nonspeech and NOISEX-92. The experimental results demonstrate that our proposed method significantly improves ASR performance, reducing the character error rate (CER) by 17.27% compared to traditional joint training methods.

Read the paper · More papers on PaperTik