Speaker identification based on sparse subspace model

Longting Xu, Zhen Yang · 2013

Mel Frequency Cepstrum Coefficient(MFCC) has been proven extremely successful for text-independent speaker identification. We address the speaker identification problem by presenting a novel Sparse Representation-Subspace algorithm. We propose to develop an overcomplete dictionary of each speaker using the Mel filterbank log energies for all the training utterances. We therefore propose to represent learned dictionary as a linear combination of all the log energies, thereby generating a naturally sparse representation, which is the novel subspace of the speaker. Besides, DCT step of MFCC is a fixed matrix, learned dictionary for different speaker is more adaptive. In the identification process, the unknown vectors of Mel filterbank log energies coefficients are projected into each subspace to decide the matching speaker. Experiments have been conducted on the speech database in our anechoic chamber, and a comparison with MFCC based speaker identification algorithms yields a favorable performance index for the proposed algorithm. Different sparsity and dictionary size shows different results.

Read the paper · More papers on PaperTik