End-to-End Model Based on RNN-T for Kazakh Speech Recognition

Orken J. Mamyrbayev, Дина Оралбекова, Aizat Kydyrbekova, Tolganay Turdalykyzy, Akbayan Bekarystankyzy · 2021

Automatic speech recognition is a rapidly developing area in machine learning. The most popular speech recognition systems today are end-to-end systems, especially those models that directly output a sequence of words taking into account the input sound in real time, which are online end-to-end models. Stream speech recognition allows to transfer the audio stream to speech-to-text conversion and get the results of stream speech recognition in real time as the audio is processed. This article discusses and implements a popular RNN-T-based model for recognizing Kazakh speech. The analysis of works related to recognition of Kazakh speech based on the CTC model is also given. The findings demonstrated that an RNN-T-based model can work well without additional components, like a language model and showed the best outcome on our dataset. As a result of the research, the system reached 10.6% CER, which is the best indicator among other end-to-end systems for recognizing Kazakh speech.

Read the paper · More papers on PaperTik