Speaker independent isolated words recognition system for Chhattisgarhi dialect
Narendra Digambar Londhe, Ghanahshyam B. Kshirsagar · 2017
Language is the main important media of communication for human beings. Automatic speech recognition (ASR) or computer speech recognition is the method or technology developed to extract, recognize and translate the speech characteristics spoken by human into text by smart computerized devises. In this paper, we have developed speaker independent ASR for a rare and geographically important Indian dialect `Chhattisgarhi'. For recognition and matching of each utterance spoken by people, we have extracted speech characteristics by using Mel frequency cepstral coefficient (MFCC) technique. The Machine Learning algorithms have been implemented on the MFCC features extracted from the self-collected chhattisgarhi speech dataset which consist of 19000 isolated words (95 words * 200 speakers). The 10-fold cross validation technique has been implemented to improve the performance of the machine learning paradigms. The designed algorithm provides 99.84% and 94.25% of accuracy using ANN and multiclass SVM with k-fold cross validation respectively. The performances of the designed machine learning algorithms have been numerically validated based on the accuracy, sensitivity and specificity.