Speech recognition for resource deficient languages using frugal speech corpus
Imran Ahmed, Sunil Kumar Kopparapu · 2012
Building speech recognition application for resource deficient languages is a challenge because of the unavailability of a speech corpus. Speech corpus is a central element for training the acoustic models used in a speech recognition engine. Constructing a speech corpus for a language is an expensive, time consuming and laborious process. This paper addresses a mechanism to develop an inexpensive speech corpus, for resource deficient languages Indian English and Hindi, by exploiting existing collections of online speech data to build a frugal speech corpus. For the purpose of demonstration we use online audio news archives to build a frugal speech corpus. We then use this speech corpus to train acoustic models and evaluate the performance of speech recognition on Indian English and Hindi speech.