Hybrid architectures for complex phonetic features classification: a unified approach

Sid‐Ahmed Selouani, Douglas D. O’Shaughnessy · 2002

This paper examines how to exploit the advantages of a hybrid approach in order to overcome the drawbacks of classic automatic speech recognition (ASR) systems faced with complex phonetic features. The key idea consists of 'boosting' the capacity of a baseline ASR system to identify features as subtle as emphasis, gemination or relevant vowel lengthening. The 'booster' part is composed of a mixture of time delay neural networks (TDNNs) using an autoregressive version of the backpropagation algorithm. We choose to carry out trials on the Arabic language, which is characterized by the presence of complex features. We use three baseline systems: hidden Markov models (HMM), optimized version of learning vector quantization algorithm (O/sup 2/LVQ1) and classical K-nearest neighbors' classifier (KNN). The reported results showed clearly the effectiveness of the approach since the three hybrid systems (HMM/TDNN, O/sup 2/LVQ1/TDNN, KNN/TDNN) perform significantly better than their corresponding baseline systems.

Read the paper · More papers on PaperTik