Speaker Identification based on MFSC voice feature extraction using Transformer
Liao Bao, Yi Zuo · 2023
Speaker identification is a type of biometric authentication technology. It can automatically identify the speaker's identity based on voice parameters. The core technology of speaker identification is to extract voice features that can best reflect the speaker's personality characteristics from the collected speech samples, and train models based on these features identify the speaker, recognize voiceprint, and so on. In the research field of speaker identification, voiceprint feature extraction briefly determines the accuracy of the speaker identification model. Among numerous voiceprint features, Mel Frequency Cepstral Coefficients (MFCC) are widely used in voiceprint identification systems due to the excellent performance of Mel filters. However, several studies revealed that MFCC features are not completely correlated globally, and only a few feature vectors are sufficient to represent most of the information in the signals. To address this limitation, we propose a new spectral representation of compressed speech, which is named as Mel Frequency Spectral Coefficients (MFSC). In MFSC, we eliminate discrete cosine transform (DCT). In the experiments, MFCC is used as the comparative feature, and end-to-end neural networks of bidirectional GRU, bidirectional LSTM, and Transformer are used as the identification models. According to 921 voice data from the LibriSpeech database, experiments have shown that the MFSC model using Transformer has better testing accuracy than MFCC models, and the error rate is reduced from 0.090 to 0.079.