Channel normalisation by using rasta filtering and the dynamic cepstrum for automatic speech recognition over the phone
Péter Boda, J.M. de Veth, Lou Boves · Radboud Repository (Radboud University) · 1996
Human auditory perception is perfectly capable to deal with time-invariant linear filter effects, such as those introduced by telephone handsets and telephone channels.We compared two different schemes for modeling human auditory time-frequency masking: RASTA filtering and the dynamic cepstrum represen tation (DCR).We used a small set of context-independent phone hidden Markov models for a recognition task o f connected digit strings over the telephone.We found that RASTA filtering out performed the Gaussian DCR approach, despite the fact that RASTA represents a more crude approximation of human for ward masking.Our results may be influenced by the choice of the mel-frequency cepstral representation that we used.The superiour performance of the RASTA technique may also be ex plained by the fact that the frequency response of the RASTA filter is better matched to the region of modulation frequencies where human auditory perception is most sensitive.