Speech database and text corpora for Malayalam language automatic speech recognition technology
Cini Kurian · 2016
Speech corpus is the backbone of an Automatic speech Recognition system. This paper presents the development of speech corpora for different Malayalam speech recognition tasks. Pronunciation dictionary and Transcription file which are the other two essential resources for building a speech recognizer are also being created. Speech corpus of about 18 hours has been collected for different recognition tasks.