Efficient Vector Code-book Generation using K-means and Linde-Buzo-Gray (LBG) Algorithm for Bengali Voice Recognition
Sadi Mohammud Syfullah, Zareen Binte Zakaria, Md Palash Uddin, Md. Fazle Rabbi, Masud Ibn Afjal, Adiba Mahjabin Nitu · 2018
Voice Recognition, an interdisciplinary field of speech processing and natural language processing, is an emerging research area for developing various real-life applications. Many techniques have been developed for recognizing voice in different natural languages. But, a few attempts have been made for Bengali voice because of its some multifaceted pronunciation nature. This paper presents a model for Bengali speech recognition which converts a Bengali voice character into its corresponding text using Mel Frequency Cepstral Coefficient (MFCC) with a presented modified procedure named K-meansLBG to generate an effective code-book. The model takes voice from user and the vector quantization (VQ) based artificial neural network (ANN) maps the Bengali voice to text notation using the generated code-book by K-meansLBG, a combination of K-means and Linde–Buzo–Gray algorithms. The model has successfully converted any single Bengali voice character into its equivalent text. The experimental outcome illustrates that the presented model produces satisfactory recognition accuracy (81.61%) on the dataset containing 900 examples of 45 different Bengali voice characters.