Far-field speech recognition Model based on Improved Conformer
Wangjun Ding, Chundong Xu, Changsong Lei, Fengpei Ge · 2024
The background noise in the far-field environment will seriously interfere with the speech signal, and it is difficult to directly apply the general speech recognition model to the far-field speech recognition. We propose an end-to-end model based on the Conv1D-Conformer architecture for far-field noisy speech recognition. The model uses an improved convolution module to replace the feed-forward neural network module, so that the encoder pays more attention to the local features that are easily disturbed by background noise. When the improved Conformer model deals with the FBANK features of far-field noisy speech data, it can extract high-dimensional features that are more consistent with the characteristics of far-field noisy speech. Connection Time Series Classification (CTC) combined with transfer learning strategy was used to assist training, which accelerated the convergence speed in the process of model training and reduced the complexity of model training. The experimental results show that the improved Conformer speech recognition model collocation transfer learning proposed in this paper achieves a relative reduction of 10.4% and 20.0% in word error rate (WER) on the far-field speech dataset CHiME-6 development set and test set, respectively.