Trace-segmentation of isolated utterances for speech recognition
Euvaldo F. Cabral, G.D. Tattersall · 2002
Trace-segmentation (sometimes called variable frame rate coding) is a method for nonlinear time-normalization of a sequence of speech representation frames prior to recognition of the sequence. Numerous attempts to perform speech recognition using trace-segmentation have been made in the past but these attempts have failed to provide the same performance as DTW or HMM recognition. The reason for this failure may be due to the use of inappropriate distance metrics to perform the segmentation or the use of an inappropriate spatial sampling interval along the trace. This paper describes an investigation into these problems, in which the appropriate Nyquist sample rate of the spatial trace is determined by analyzing the frequency of the temporal variation of the speech frames. It is also shown that separate segmentation of the trajectory described by each individual coefficient in the speech frame leads to much improved recognition which exceeds the performance provided by DTW recognition of the same database.