Dual-Branch Speech Enhancement Network for Noise-Robust Automatic Speech Recognition
Haobin Jiang, Tianlei Wang, Dinghan Hu, Jiuwen Cao · 2024
Automatic speech recognition (ASR) has achieved remarkable successes thanks to the end-to-end deep neural networks, but it is still challenging in the noisy and reverberation environments. The joint training of front-end speech enhancement (SE) and speech recognition system becomes a popular solution to the noise-robust ASR. However, the speech distortion problem arisen due to excessive information suppression by the SE module. To address this issue, in this paper, a novel dual-branch SE (DBSE) module is proposed as the front-end of the ASR system for joint training. Particularly, the two branches extract the clean speeches using different ways: one branch directly extracts clean speeches and the other branch utilizes spectral subtraction method. The final denoising speech signals are obtained by combining the outputs from both branches. In this way, the oversuppressed information can be compensated from each other. The two-stage joint training strategy is adopted for the noise-robust ASR model where the proposed DBSE is first pretrained by multitask reconstruction loss, and then the DBSE and ASR model is jointly trained using the speech recognition based loss function. Comparisons with several state-of-the-art ASR algorithms on benchmark dataset are conducted, and the results demonstrate the superior performance of our proposed algorithm.