DWT-based classification of acoustic-phonetic classes and phonetic units
Gernot Kubin, Van Tuan Pham · 2004
Abstract In this paper, we describe a new algorithm based on the discrete wavelet transform (DWT) which uses a multi-threshold decision model (MTD model) to detect acoustic and phonetic classes (based on 10ms speech signal segments). The best thresholds of the model are found by using experimental pattern classification. Then a unit level interpolation technique is combined with the MTD model to classify phonetic units (based on sequences of 10ms segments). The results of the classifiers are compared and jointly adjusted by an interactive scheme (IS) in order to improve the performance of the algorithm. The algorithm is tested with the TIMIT database and compared with the SUB-CRA-based algorithm and other algorithms to demonstrate its effectiveness. 1 {( )}Introduction The speech classification problem plays an important role in many speech processing algorithms and applications. Its solution can improve performance of concatenative speech synthesis by selecting proper smoothing strategies at concatenation points. Some speech coding systems use phonetic classification to determine the optimal bit allocation for each speech segment. The discrimination of different speech classes has been studied in many articles since the 1980’s. Most of them exploit statistical features of the speech signal [1]–[2] such as relative energy level, zero crossing rate, etc., to decide about voiced, unvoiced, silence or mixed excitation. However, these approaches only achieve good accuracy if using a large number of parameters. In recent years, wavelet analysis attracted researchers in speech applications. Some algorithms based on DWT have been developed to classify speech at the phonetic level [4]-[7]. The combination of wavelet parameters with statistical parameters opens a way to increase classifier performance. Especially, this combination can be exploited to solve more difficult tasks such as classification of phonetic units [3],[8].