Bangla Speech Recognition for Voice Search

Jillur Rahman Saurav, Shakhawat Amin, Shafkat Kibria, M. Shahidur Rahman · 2018

In this work, different Gaussian Mixture Model-Hidden Markov Model(GMM-HMM) based and Deep Neural Network (DNN-HMM) based models have been analyzed for speech recognition in Bangla language to build a voice search module for search engine pipilika1. A small corpus of 9 hours of speech recordings from 49 different speakers was prepared for this work consisting of a vocabulary of 500 unique words. The lowest Word Error Rate(WER) for (GMM-HMM) based model was 3.96% and for (DNN-HMM) based model was 5.30%. To our best knowledge, this is the lowest WER for Bangla speech recognition for such vocabulary size.

Read the paper · More papers on PaperTik