Syllable-level desynchronisation of phonetic features for speech recognition

Katrin Kirchhoff · 2002

Describes a novel approach to speech recognition which is based on phonetic features as basic recognition units and the delayed synchronisation of these features within a higher-level prosodic domain, viz. the syllable. The object of this approach is to avoid a rigid segmentation of the speech signal as it is usually carried out by standard segment-based recognition systems. The architectural setup of the system is described, as well as evaluation tests carried out on a medium-sized corpus of spontaneous speech (German). Syllable and phoneme recognition results are given and compared to recognition rates obtained by a standard triphone-based HMM recogniser trained and tested on the same data set.

Read the paper · More papers on PaperTik