A multi-spectral data-fusion approach to speaker recognition
J. E. Higgins, R.I. Damper, Chris J. Harris · ePrints Soton (University of Southampton) · 1999
Abstract This paper describes a multi-spectral, multisource approach to the important problem of speaker identification. The wideband speech signal is filtered into several sub-bands and the output time trajectory of each is individually modeled by linear prediction cepstral coefficients. These individual models are then matched against reference data and the scores combined using the sum rule of information fusion, before using a k-nearest-neighbor rule to decide the identified speaker. Multi-spectral processing is shown to deliver performance improvements over wideband recognition. The optimal number of filters is found to be 16. These results are interpreted in light of the hypothesis that the multi-spectral approach solves the bias/variance dilemma of commonly manifest in systems that are trained on example data.