Speech Recognition Based Web Browsing
Vinitkumar Jayaprakash Dongre, Sanjeev Ghosh, Ashwini Katkar · 2013
The technology of voice browsing is rapidly evolving these days. Listening and speaking are the natural modes of communication and information gathering. As a result all are now heading towards a more voice based approach of browsing rather than operating on textual mode. The results of a case study carried out while developing an automatic speech recognition system for web browsing are presented in this paper. A generalized coding is done to make the system compatible for „n' number of samples without any change in basic coding. Speech data is collected from independent speakers and pre-processed to extract the features needed in this research. Feature extraction is carried out using Mel Frequency Cestrum Coefficient (MFCC) technique. In [1], it was demonstrated that MFCC outperforms than other feature extraction techniques. After the training session, the acoustic vectors extracted from input speech of a speaker provide a set of training vectors. The centroid based neural network Adaptive Resonance Theory (CNNART) approach is used for mapping vectors from a large vector space to a finite number of regions in that space. For comparison purpose, the distance between each test codeword and each codeword in the master codebook is computed. The difference is used to make recognition decision. The prototype can recognize the word as well as sentences by concatenating the words stored in the database to form a sentence. The recognition accuracy of the system is 85% in speaker dependent environment while 70% in speaker independent environment. Also, system provides 70% accuracy for sentence recognition while for isolated word, recognition accuracy is 80%.