The Research of Noise-Robust Speech Recognition Based on Frequency Warping Wavelet
Xueying Zhang, Wenjun Meng · 2007
The later conculsions can be obtained from the experimental results of the fourth section. 1. The three parts of speech recognition are conjunct one another and exist the relation restricted among themselves. Bark wavelet was used in improving the feature of ZCPA and MFCC the latter effect is obviously better than the former.It illustrates that Bark wavelet and the speech character described by MFCC feature are more closer than Bark wavelet and the speech character described by ZCPA feature. The fact is also as such. Bark wavelet is constructed directly according to the hearing perception of human ear, and MFCC is the cepstrum coefficents on the basis of Mel frequency. While Mel frequency is just the hearing frequency of human ear. Though the frequency bins of ZCPA are divided according to the hearing perception, the zero-crossing rate and peak amplitude are time-domain parameters which are transformed nonlinearly mapping to frequency bin. This kind of nonlinear transform may affect the consistency of ZCPA and hearing frequency, that results in decreasing in function. If the selection of training or recognition network is different, they have different effect on the results. Furthermore, the function of recogntion network has direct relationship with front-end filter and feature extracted. This point can be seen from the experimental results of combination mode1 FIR+ZCPA+HMM and mode 2 FIR+ZCPA+WNN . Comparing the two modes, the former two parts are same and the third part is different from using HMM or WNN the results obtained have much more different.The wavelet neural network has bright foreground for speech recognition.Its training speed is fast, which is good for implementation in real time. Further,it has also good recognition rates under no noise or noise environment and the number of recogintion words is larger. The paper researched some kinds combination modes aiming to the three parts of speech recognition system in Fig. 1. For other combination modes, such as Bark+MFCC+WNN Bark+ZCPA+WNN and so on, we will research them in later work. Which of combination ever is optimal? This needs considering practical application case. We hope the research can be refered by interesting researcher and get to the purpose of communication mutually and progress.