Introduction to Automatic Speech Recognition: Template Matching
John Holmes, Wendy J. Holmes · 2002
This chapter describes typical methods that were developed for spoken word recognition during the 1970s. A well-established approach to Automatic Speech Recognition is to store in the machine example acoustic patterns for all the words to be recognized, usually spoken by the person who will subsequently use the machine. In calculating a distance between two words it is usual to derive a short-term distance that is local to corresponding parts of the words, and to integrate this distance over the entire word duration. A consequence of removing the effect of the fundamental frequency and of using filters at least as wide as critical bands is to reduce the amount of information needed to describe a word pattern to much less than is needed for the waveform. A suitable distance metric for use with a filter bank is the sum of the squared differences between the logarithms of power levels in corresponding channels.