Comparison of models using time-frequency features for speech classification
Tuan Van Pham, Gernot Kubin · 2006
This paper reviews and evaluates two different ways of speech classification using multidimensional features derived from the Dyadic Wavelet Transform (DyWT). The first method is based on the Multi-Threshold Decision Model, while the second method is based on the training of Feedforward Neural Networks. The extracted features are compared against adaptive thresholds of the models to classify the input speech signal frames in terms of acoustic classes and phonetic groups. The methods are tested and assessed with the TIMIT database in counting gender dependency. The paper presents the procedures of building the speech classifiers, shows the performance comparison of the two algorithms with other algorithms, and discusses their advantages as well as limitations. Automatic speech classification is crucial for many speech processing methods and in various speech applications. Voice activity detection which is employed by most speech applications is a direct application of speech classification. In automatic speech recognition, there is a need for a phonetic classifier to improve the performance of endpoint detection in order to increase the recognition rate. The performance of concatenative speech synthesis may be improved by selecting proper smoothing strategies at concatenation points. Some speech coding systems use speech classification to determine the optimal bit allocation for every different speech frame. In internet telephony applications, the adaptive loss concealment algorithm uses the voiced/unvoiced detector at the sender. This helps the receiver to conceal the loss of information based on the similarity between the lost segments and the adjacent segments. Besides, the phonetic alignment of huge databases can be performed faster by applying the phonetic classifier as a pre-classification step. The speech classification task has been studied in many articles by a variety of methods since the 1980's. In principle, the classification is done by relying on different types of feature vectors which are extracted from the input speech frames. These features can be derived by three approaches: • The first approach works in the time domain and uses statistical measurements. The common features are zero crossing rate, relative energy level, autocorrelation coef- ficients, etc. (1)-(3). This approach only achieves good accuracy if using a large number of parameters.