Adapting automatic speech recognition methods to speech perception: A hidden semi-Markov model of listener’s categorization of a VC(C)V continuum

Terrance M. Nearey · The Journal of the Acoustical Society of America · 2004

This study reanalyzes the categorization of 144 synthetic VC(C)V stimuli by 13 listeners. Medial silent gap duration and frequencies of pre- and postgap F2−F3 transitions were independently varied to cover the response set / aba, abda, ab♯ba, ada, adba, ad♯da/. A simple logistic model was shown to work very well for the complex response patterns [T. Nearey and R. Smits, J. Acoust. Soc. Am. 24, 111, 2434 (2002)]. The former analysis required a static ‘‘spoon-fed’’ description of the variable duration stimuli in terms of synthesis parameters. A new analysis will be presented using a hidden semi-Markov model (HSMM), which is a fully automatic dynamic pattern recognizer. The HSMM includes explicit state durations and is fitted to perceptual data by optimization methods allied to the maximum mutual information (MMI) approach from the automatic speech recognition literature. Given only the waveforms of the stimuli as input, the trained HSMM provides an extremely good fit to the perceptual data. The fully automatic HSMM performs better than the spoon-fed static logistic model, with rms error values of 4.8 versus 5.9 percent, respectively. The flexibility and generality of the HSMM/MMI framework for perception models will be sketched. [Work supported by SSHRC.]

Read the paper · More papers on PaperTik