Based on the Conformer End-to-End Chinese Speech Recognition Method
Haiwen Feng, Yunxiao Shi · 2024
Aiming at the problem that the acoustic input network of Conformer encoder is insufficient for Fbank speech feature extraction, an end-to-end speech recognition modeling method of DM-Conformer is proposed. Firstly, the Dilated convolution is utilized in the downsampling module to increase the receptive field, which captures more speech feature information without adding extra model parameters, while ensuring that the height and width of the original input feature map remain unchanged. Then, the Mish activation function is introduced to alleviate the problem of slow convergence of the ReLU activation function due to its negative value. Finally, experiments are carried out on the public dataset AISHELL-1, and the speech recognition model with the improved downsampling module proposed in this paper reduces the word error rate by 5.6% compared with the Conformer baseline model on the test set, and the model finally achieves a word error rate of 6.15%.