Stochastic perceptual models of speech
N. Morgan, H. Bourland, Sarah Stein Greenberg, Hynek Heřmanský, Su-Lin Wu · 2002
We have developed a statistical model of speech (based on auditory perceptual criteria) that avoids a number of current constraining assumptions for statistical speech recognition systems, particularly the model of speech as a sequence of stationary segments consisting of uncorrelated acoustic vectors. We further wish to focus statistical modeling power on perceptually-dominant and information-rich portions of the speech signal, which may also be the parts of the speech signal with a better chance to withstand adverse acoustical conditions. We describe some of the theory, along with some preliminary experiments. These experiments suggest that the regions of acoustic signal containing significant spectral change are critical to the recognition of continuous speech.