Equalizing sub-band error rates in speaker recognition

Roland Auckenthaler, John S. D. Mason · 1997

Recent work on ASR by [1] [2] shows that band splitting gives recognition accuracy comparable with the conventional fullband. Sub-bands have different bandwidth spaced on a mel scale. Interestingly in the contex of speaker recognition improved accuracy has been reported in the case of a full-band approach using a linear scale. We demonstrate that both of these scales are likely to be suboptimum in the context of band splitting. We then describe, how sub-band error profiles can lead to a new scale, which is between a linear and a mel spacing, giving both an equalised sub-band error profile and an improved overall recognition accuracy. 1. INTRODUCTION This paper is concerned with splitting the conventional acoustic representation into sub-band units and processing these separately, with recombination at the decision stage. This idea has been investigated recently in the context of speech recognition [1] [2]; here our task is the complementary one of speaker recognition. Potential bene...

Read the paper · More papers on PaperTik