Power-Normalized Cepstral Coefficients (PNCC) for Punjabi Automatic Speech Recognition using phone based modelling in HTK
Arshpreet Kaur, Amitoj Singh · 2016
Punjabi is a regional language having variant pronunciations and ranging tones. Therefore developing a robust speech recognition system for such language is today's vital need. Conventional research practices lay stress on using basic extraction techniques like Frequency Cepstral Coefficients (MFCC), Perceptual Linear Prediction (PLP) etc. This papers subject the application of a novel Feature Extraction Technique called Power-Normalized Cepstral Coefficients (PNCC) on connected words in Punjabi Automatic Speech Recognition. This particular technique is an extension to conventional MFCC as MFCC uses mel filter bank and traditional log non linearity on other end the proposed technique uses gammatone filter bank and power law non linearity in acoustic modelling phase. As phoneme sounds (tones) are focused, phone based modelling has been enrolled using 34 phones for training 158 words. The phones are used to break each phoneme word into frames based on the sound produced. Training was done in noise-free environment using 16 speakers and noisy environments with 12 speakers. HTK (Hidden Markov Model toolkit) is used to build acoustic model hatching 8 HMM models for training. Each speaker sheered the speech corpus for number of times to increase the redundancy of the system hence, blooming the recognition rate. Performance analysis was done in noisy-free and noisy environments indulging 3 speakers giving 83.05% accuracy in noisy-free and 71.92% accuracy in noisy environment.