Speaker identification based on PCC feature vector

Peichao He, Yi Zuo, Tieshan Li, C. L. Philip Chen, He Ma, Junxia Liu · 2019

In the research field of speaker identification, many extraction methods of speech feature have been investigated for several decades. However, several new features such as i-vector and d-vector were proposed in recent years. Mel Frequency Cepstral Coefficient (MFCC) is still wildly used in current speaker identification systems accounting for its high performance. Based on the similar generation approach of MFCC, this article proposes a novel feature extraction way based on Pearson correlation coefficient (PCC). Firstly, we also use inverse discrete cosine transform (IDCT) cepstrum coefficient as the initial speech inputs. Secondly, we employ a hierarchical clustering analysis based on PCC to merge the IDCT cepstrum coefficient until the dimension of speech inputs is reduced to 14. Finally, we output this 14-dimensional vector as speech feature named r-vector. In the experiments, Gaussian Mixture Model (GMM) was applied to compare the performance of r-vector with other speech features. According to the 630 people voice data in TIMIT database, the results of experiments claim that the r-vector could obtain higher recognition accuracy in speaker identification.

Read the paper · More papers on PaperTik