Research on Speaker Identification Models Based on CNN and Additive Angular Margin Loss
Xi Xuan, Runping Han, Baiyong Ding · 2021 2nd International Conference on Electronics, Communications and Information Technology (CECIT) · 2021
In view of the traditional speaker identification methods which usually have poor generalization ability and high computational complexity, three text-independent closed-set speaker identification models named Res-SD, Res-SArc and Rep-SArc respectively are designed based on the Red-Song dataset constructed by us. Res-SD is trained by using the traditional cross-entropy loss, while Res-SArc and Rep-SArc are trained by using additive angular margin loss that optimizes the feature embedding to realize higher intra-class similarity and higher inter-class diversity. In terms of the number of model parameters, Rep-SArc accounts for only 1/3 of Res-SArc. Test experiments for evaluating three models are conducted on the Red-Song dataset's test set. In terms of identification accuracy, Rep-SArc and Res-SArc are more competitive among three models, and their identification accuracies reach 97.90% and 97.19%, respectively. The experimental results validate that these three models are effective for the text-independent closed-set speaker identification tasks.