Far-Field Speech Recognition Based on Complex-Valued Neural Networks and Inter-Frame Similarity Difference Method

Yifan Guo, Yifan Chen, Gaofeng Cheng, Pengyuan Zhang, Yonghong Yan · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) · 2021

Far-field automatic speech recognition (ASR) is a challenging task due to the background noise and reverberation. To address this issue, we introduce a novel end-to-end multi-channel far-field ASR architecture. First, we use a complex-valued CNN based architecture designed for speech tasks as a neural beamformer. Second, we propose an auxiliary mod-ule called absolute position regression module (APRM) with a position prediction loss to help the neural beamformer be better aware of the corresponding frequencies of each input time-frequency (T-F) bin. Third, inspired by the short-term stationarity of human speech, we propose an approach called the Inter-Frame Similarity Difference (IFSD) method to au-tomatically select useful channels as the inputs of the ASR backend from the outputs of the neural beamformer. We also implement a complex-valued attention module for the output channels of the neural beamformer to utilize each other's in-formation, thereby preventing the final outputs from information loss. With the above innovations, our proposed model achieves 9.7% and 11.1% relative WER reductions over a DNN-MVDR baseline on the CHiME4 dataset and a dataset simulated using the Librispeech corpus.

Read the paper · More papers on PaperTik