Design and development of a large vocabulary, continuous speech recognition system for Tamil

A. Madhavaraj, A. G. Ramakrishnan · 2017

This paper presents our work on building a large vocabulary continuous speech recognition system for Tamil using deep neural networks (DNN). Well known techniques, namely, maximum likelihood linear transformation and speaker-adaptive training have been used to build our final deep neural network based speech recognition system. We have used 6.5 hours of Tamil speech recorded from 30 speakers covering a vocabulary of 13,026 words, of which 4.5 hours of data was used for training, 1 hour of data for testing and 1 hour of data for cross-validation. Two independent recognition systems were built, one for phone recognition (PR) and the other for continuous speech recognition (CSR) and they achieve phone error rate of 24.9% and word error rate of 3.5%, respectively. DNN-based triphone acoustic model shows an absolute improvement of about 1% and 23% over the monophone acoustic model for CSR and PR, respectively.

Read the paper · More papers on PaperTik