Self-attention Networks for Speaker Identification with Negative-Focused Triplet Loss
Jianye Mo, Xu Li · Journal of Physics Conference Series · 2020
Abstract In this paper, a model based on self-attention for text-independent speaker identification task is proposed. Traditional CNNs are expert at capturing local features but are inefficient when capturing long-range dependencies. The self-attention helps to capture long-range dependency more efficiently and requires less parameters. Besides, based on Triplet Loss and inspired by Weighted Triplet Loss, we propose a novel loss function for speaker identification task, named Negative-Focused Triplet Loss, which makes the training process more efficient and effective. To evaluate the performance of our model, the experiments are conducted on Voxceleb dataset, and we achieve Top-1 accuracy of 90.3% and Top-5 accuracy of 96.9%, which are competitive with the previous state-of-the-art performance. Without any hard sample selecting operation, we achieve better result than the baseline method that uses Cluster-Range Loss or Triplet Loss, which demonstrates the high efficiency of our proposed approach.