Speech-to-music ratio estimation using wavelets and hidden Markov models
Brett Y. Smolenski · The Journal of the Acoustical Society of America · 2007
In this paper, a system capable of estimating the long-term, on the order of an utterance, speech-to-music energy ratio (SMR) is developed. The approach uses frame-based Teager energy values at the output of a seven-band wavelet decomposition as features. The features are then modeled using a two-state hidden Markov model (HMM), with each state producing observations having a 64-component Gaussian mixture. Using this approach, estimation of the long-term SMR with a standard error of less then 5% was obtained. In addition, accurate classification of music and speech with these models was still possible when both music and speech were simultaneously present, assuming one had a higher energy than the other.