A study of multilingual speech features: perceptive scalogram based on wavelet analysis
Md Ismail Hossain, J. Liu, R. Lee · 2003
In recent years adopted speech feature for automatic speech recognition system is degraded by the additive noise when the system is being used in real environments. In this respect a lot of efforts have been devoting to find the robust features. In this paper, we propose the use of perceptive scalogram, which is a kind of time-scale distribution of signal to approximate perception of speech. It provides a three-dimensional representation of speech signal spectra, with time represented on the x-axis and frequency on the y-axis. Gray scales are used to represent the third dimension, either the amplitudes (for scalograms) or the phase (for phase-shift diagrams), of the wavelet transform data. It takes considerations that speech and background noise has different perceptions on human beings, and speech signals are non-stationary. It is a time-scale distribution of signal, and uses wavelets as its analysis tool, which makes a good approximation of some hearing properties. Computer simulations show that the proposed perceptive scalogram is very robust to additive white Gaussian noise for different languages (e.g. English, Putonghua, Cantonese).