Speech coding for fast speech commands recognition

Abel Herrera, Arturo Gardida · The Journal of the Acoustical Society of America · 2002

Isolated speech recognition could be more appropriate than continuous speech recognition systems for some word command applications. In this paper, some speech coders are explored that can be useful for fast word command recognition using low memory storage. First, decimated critical bands vectors by frame are used as coding features. The accuracy is around 94% for a set of ten commands; they are used at first step. After that, using a reduced DTW classification scheme, the accuracy for a small subset of possible candidates is above 99%. Another coding scheme is using the Karhunen–Loeve Transform (KLT) combined with multisection vector quantization. The KLT is not a fast transform; however, used at the subword level it gives a fast recognizer with accuracy around 99%. The classifiers in both cases are fast, around a second, and require low memory storage.

Read the paper · More papers on PaperTik