Robust voiced/unvoiced speech classification using empirical mode decomposition and periodic correlation model
Md. Khademul Islam Molla, Keikichi Hirose, Nobuaki Minematsu · 2008
Abstract This paper presents a method of voiced/unvoiced (V/Uv) classification of noisy speech signals. Empirical mode decomposition (EMD), a newly developed tool to analyze nonlinear and non-stationary signals is used to filter the additive noise with the speech signal. The normalized autocorrelation of the filtered speech signal is computed to enhance the periodicity if any. It is considered that the voiced speech signal is periodically correlated and the unvoiced signal is not. A statistical model of determining periodic correlation is used to differentiate voiced and unvoiced speech with low SNR. The experimental results show that the use of EMD improves the classification performance and the overall efficiency is noticeable as compared to other existing algorithms. Index Terms : empirical mode decomposition, normalized autocorrelation, periodic correlation, voiced/unvoiced speech 1. Introduction Reliable classification of short time speech signal into voiced and unvoiced is a crucial preprocessing step in many speech processing applications and is essential in most analysis and synthesis system. For example: different strategy could be adopted for voiced and unvoiced parts in speech enhancement using spectral subtraction. The essence of classification is to determine whether the speech production system involves the vibration of the vocal cords [1]. The discrimination problem is an important one and has been worked on extensively during the last three decades [2]. The discrimination can effectively be performed using a single feature or parameter which is closely associated with the voicing and non-voicing activities of speech signal. Many algorithms have been reported for solving the detection problem [3] – [7]. In [3], Gaussian mixture model with cepstrum coefficients features is proposed for robust V/Uv classification. A higher order statistics (HOS) based method is proposed in [4] for V/Uv detection and pitch estimation simultaneously. The matching pursuit algorithm is used in [5] with Gabor decomposition. The wavelet transform is proposed in pitch and V/Uv detection in [6]. A statistical model applied in autocorrelation domain is also reported in [7]. In most of the existing algorithms are not so much noise robust and also the intensive threshold and training data are required for classification. Such requirements are troublesome for the use in application domain. The proposed method is noise robust and based on the statistical model for periodicity detection in speech signal without any training requirement. To reduce the effect of noise on speech signal, a data adaptive time domain filtering is proposed using newly developed empirical mode decomposition method [8]. Although speech signal is non-stationary in nature, Fourier based frequency domain filtering assumes that it is piecewise stationary. The speech decomposition is performed by fitting some predefined bases without satisfying its non-stationary nature. Whereas, EMD based approach decompose the speech signal as non-stationary time series and hence better performance in noise filtering. A method for determining whether an observed time series contains a periodically correlated sequence is employed here. It is based on the statistical tests for the coherence between spectral components for the presence of a periodically correlated covariance structure in a time series [9]. The autocorrelation function (ACF) makes the periodicity more prominent if any. The proposed periodic correlation model is applied in the autocorrelation domain rather than original time domain of the speech signal. The periodicity detection method is implemented in spectral domain to classify the speech segment into voiced or unvoiced one based on that it contains periodic correlated sequence or not respectively.