Speaker Recognition with ResNet and VGG Networks

Maroš Jakubec, Eva Lieskovská, Roman Jarina · 2021

Various deep neural network (DNN) topologies have been recently introduced for automatic speaker recognition (SR) tasks, which tend to have ever deeper architectures. Several specific architectures have been proposed to tackle the well-known vanishing gradient issue in DNN training. We present results of experiments on text-independent SR in the wild with two deep Convolutional DNN architectures: ResNet and VGG. The results on the VoxCeleb1 benchmark database demonstrate the superiority of the ResNet solutions, in comparison with VGG and standard i-vector approaches, for speaker identification and verification in realistic acoustic conditions.

Read the paper · More papers on PaperTik