Voiced-unvoiced decision without pitch detection
Bishnu Saroop Atal, L. R. Rabiner · The Journal of the Acoustical Society of America · 1975
In speech analysis, the voiced-unvoiced decision is usually performed in conjunction with pitch analysis. The linking of voiced-unvoiced decision to pitch analysis not only results in unnecessary complexity but makes it difficult to classify short speech segments. We present a method of classifying a short speech segment, typically 10 msec in duration, into three classes: voiced, unvoiced, and silence. For each of the three classes, a nonEuclidean distance is computed from the measurements made on the speech segment to be classified and the segment is assigned to the class with the minimum distance. The distance metric used is given by (x − μi)t Wi−1(x − μi) where x is the vector representing the speech measurements, μi is the mean vector, and Wi is the covariance matrix for the ith class. The speech parameters are the zero-crossing rate, the speech energy, the correlation between adjacent speech samples, the first predictor coefficient, and the prediction-error energy. A training set of data is used to obtain the mean vectors and the covariance matrices for the three classes. A simple median-smoothing algorithm is used to eliminate isolated errors. The method has been found to provide reliable performance for both speech synthesis and speech-recognition applications.