Development of a Bangla Speech to Text Conversion System Using Deep Learning
Srijoni Saha, Asaduzzaman Asaduzzaman · 2021
Bangla Speech-To-Text (STT) conversion is a technology that provides a means of converting spoken Bangla language to a written Bangla text form. The standard of speech recognition in different languages is rising step by step however Bangla speech recognition has drawn exceptionally little attention. Building up an STT framework is a bulky procedure and it requires a few stages. Deep neural network-based architecture replaces the stages with neural network components which makes the task simpler and removes the dependency on hand-engineered rules. CNN-RNN networks with CTC criterion are utilized in this undertaking to construct a Bangla STT system that generates text from speech. The architecture is trained solely on speech samples and text transcripts. We used 215.53 hours of speech data set for training, which includes a wide variety of speech samples collected by people of different ages and genders. This paper shows a comparison between genuine text transcripts with produced text transcripts for a similar sound example. Comparison of results between implemented architecture and already existing Bangla STT has also been presented in this paper. The paper is concluded with a discussion about the word error rate and implementation challenges. The key contribution of this paper is to develop a gender and speaker-independent continuous speech-to-text conversion system for the Bangla language using deep learning.