A Large Vocabulary End-to-End Myanmar Automatic Speech Recognition

Hay Mar Soe Naing, Win Pa Pa · 2023

In recent years, sequence-to-sequence technology has become popular in automatic speech recognition area. This model replaces the classic complex pipeline with a single neural network architecture. This paper proposes the use of transformer- and conformer-based models on Myanmar automatic speech recognition system (UCSY-Myan-ASR). Classical hybrid long short-term memory (LSTM) and end-to-end models are presented and evaluated to improve error rates. The experiments were carried on the UCSY-82-hour speech corpus and evaluated in terms of syllable error rate (SER) and character error rate (CER). Using the Transformer approach, the best performance in the daily conversation domain reaches the SER of 9.6% and CER of 7.3%. When using the conformer model, the best performance in the news domain is 10.6% SER and 6.9% CER respectively.

Read the paper · More papers on PaperTik