An efficient combination of acoustic and supra-segmental informations in a speech recognition system
Nelly Suaudeau, Régine Andre-Obrecht · 2002
A major deficiency of a standard HMM is that both the spectral and the prosodic features are uniformly processed. To more efficiently combine the prosodic cues together with the acoustic ones, a two level HMM which separates the spectral and suprasegmental representations is defined. Namely, the incorporation of global sound durations is explored. More, to take into account the effects of speaking rate on the phonetic unit durational parameters, two durational models are proposed. The ways those models are integrated in the recognition processing are described. Experiments on a French number database show that such an explicit introduction of prosodic parameters reduces recognition errors rates by 20%.>