Speaker Recognition Based on Ensemble Neural Network
Yao Xin He, Pengfei Shuai, Peifan Jiang, Gang Tang · 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI) · 2022
To improve the accuracy of speaker recognition, this paper proposes a speaker recognition method based on an ensemble neural network. The main idea of our method is to fuse the three basic networks of EfficientNet, ResNet, and GoogleNet to build an EnsembleNet. A weighted average of three networks obtains the output of this new model. A weighted average of the three networks is used to produce the output of this new model. The secondary network, which is used to train the weighted average weight values, receives the output of the basic network as input, and the weights vary adaptively throughout the training process. First, the basic network is trained individually, and the training stops when the basic network's recognition accuracy is the highest. Next, the weights are trained, and the output of the basic network prediction is used as the input of the secondary network to train the weighted average weight value. Each basic network produces a feature vector whose length is the number of speakers, whereas the secondary network produces a vector whose length is the number of basic networks. The proposed method is validated on the ST-CMDS Chinese open source dataset and the ST-CMDS, AIshell, Aidatatang hybrid dataset, and Zhvoice dataset. The experimental results show that our method is better than the three separation models and the voting method. And the recognition accuracy rate reached 97.02%, 95.04% and 53.97%.