Frequency Attention Module for Speaker Recognition
Yujiao Wu, Zulin Fang, Jianping Dong, Gexiang Zhang · 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI) · 2022
Speaker recognition is a task of identifying user’s individuality from their speech utterances. Most of the text-independent speaker recognition models available directly input the spectrogram into the deep neural network. However, this does not take full advantage of the spectrogram features. In this paper, we propose a frequency attention module (FAM) based on the fact that the frequency axis in the spectrogram is more representative for speaker feature extraction. It aims to assigning higher weights to important frequency bins, and making the network to extract more features to improve speaker recognition performance. Furthermore, FAM is applied on the convolution output in ResBlock of ResNet34 to capture richer features of the speaker. The experimental results under self-attention pooling (SAP) and attentive-statistics pooling (ASP) show that FAM is feasible and effective. Moreover, the extended experimental results on the Free ST Chinese Mandarin Corpus dataset indicate that the proposed FAM (equal error rate (EER) is 1.59%) outperforms the no-attention ResNet34 (EER is 1.97 %) and spatial-CBAM ResNet34 (EER is 1.88%).