Speech recognition using wavelet packets, Neural Networks and Support Vector Machines
Purva Kulkarni, Saili Kulkarni, Sucheta Mulange, Aneri Dand, Alice N. Cheeran · 2014
This research article presents two different methods for extracting features for speech recognition. Based on the time-frequency, multi-resolution property of wavelet transform, the input speech signal is decomposed into various frequency channels. In the first method, the energies of the different levels obtained after applying wavelet packet decomposition instead of Discrete Fourier Transforms in the classical Mel-Frequency Cepstral Coefficients (MFCC) procedure, make the feature set. These feature sets are compared to the results from MFCC. And in the second method, a feature set is obtained by concatenating different levels, which carry significant information, obtained after wavelet packet decomposition of the signal. The feature extraction from the wavelet transform of the original signals adds more speech features from the approximation and detail components of these signals which assist in achieving higher identification rates. For feature matching Artificial Neural Networks (ANN) and Support Vector Machines (SVM) are used as classifiers. Experimental results show that the proposed methods improve the recognition rates.