Speaker Independent Speech Recognition Using Maximum Likelihood Approach for Isolated Words
Amaresh P. Kandagal, V. Udayashankara · INTERNATIONAL JOURNAL OF COMPUTER APPLICATION · 2017
Speech is an intuitive interface for man machine interaction.Minimizing word error rate is a unique challenge to develop Automatic Speech Recognition (ASR) system.Performance of this system is far from perfect.Acoustic model and language models are fundamentals to build robust ASR engine.This paper presents a stochastic procedure for developing phoneme and word level acoustic models.Acoustic features estimated by Mel Frequency Cepstral Coefficients (MFCC) with 35% of overlapping of frames for every 25 milliseconds of a signal.The paper compares and highlights the word and phoneme level acoustic model performances for Kannada language vocabulary.The performance of the system is recorded for different vocabulary sizes, and word error rate (WER) computed for phoneme and word acoustic models.The system presents accuracy of 94.78046% and 97.6% for word and phoneme acoustic models respectively for the vocabulary 90 words.In addition, 98.08% of recognition rate for the vocabulary of 70 words.