Tibetan speech recognition based on wenet
Runyu Zhe, Guanyu Li, Like Ma · 2024
In our study of speech recognition for Tibetan, we have used the end-to-end speech recognition framework Wenet.This study provides an in-depth analysis of Wenet's performance under four different decoding strategies, which include an attention mechanism-based decoder, CTC greedy search, CTC prefix bundle search, and attention rescoring methods. In addition, we also comparatively analyze the training results using two different encoders (Conformer [1] and Transformer [2]).The experimental results show that the Conformer encoder exhibits generally better performance than the Transformer encoder in training, and achieves significant results in reducing the error rate, with specific percentage reductions of 2.85%, 3.02%, 3.71%, and 3.58%, respectively. This finding not only verifies the high efficiency of the Wenet framework in handling the Tibetan speech recognition task, but also highlights its significant advantage in improving recognition accuracy.