Automatic Phonetic Transcription for read, extempore and conversation speech for an Indian language: Bengali

K E Manjunath, K. Sreenivasa Rao · 2014

In this work, we have analyzed the proposed Automatic Phonetic Transcription (APT) approach for read, extempore and conversation modes of speech for Bengali language. In our earlier work, the APT was carried out using read speech. In this paper, main focus is on deriving APT for Extempore and Conversation modes of speech in Bengali language and their analysis. This framework of deriving APT can be extended to any Indian language. The Automatic Phonetic Transcription Systems (APTS) were developed separately for read, extempore and conversation modes of speech. In this study, APT has been carried out on read, extempore and conversation modes of speech using 35, 33 and 30 phones respectively. APT has been carried out using Hidden Markov Models (HMMs) and FeedForward Neural Networks (FFNNs). Mel-frequency Cepstral Coefficients are used as features for building the models. The best obtained performance accuracies using HMMs for read, extempore and conversation modes are 41.65%, 29.20% and, 23.48% respectively. Using FFNNs, the recognition accuracies for read, extempore and conversation modes are 53.87%, 46.19% and 33.63% respectively.

Read the paper · More papers on PaperTik