Improved hindi broadcast ASR by adapting the language model and pronunciation model using a priori syntactic and morphophonemic knowledge

Preethi Jyothi, Mark Hasegawa‐Johnson · 2015

In this work, we present a new large-vocabulary, broadcast news ASR system for Hindi. Since Hindi has a largely phone-mic orthography, the pronunciation model was automatically generated from text. We experiment with several variants of this model and study the effect of incorporating word bound-ary information with these models. We also experiment with knowledge-based adaptations to the language model in Hindi, derived in an unsupervised manner, that lead to small im-provements in word error rate (WER). Our experiments were conducted on a new corpus assembled from publicly-available Hindi news broadcasts. We evaluate our techniques on an open-vocabulary task and obtain competitive WERs on an unseen test set. Index Terms: Hindi LVCSR system, Broadcast news ASR, Grapheme and phoneme-based models, Knowledge-based language-model adaptation 1.

Read the paper · More papers on PaperTik