Res2Net based Text Independent Speaker recognition system

Mukund Kumar Roy, Ushaben Keshwala · 2022 12th International Conference on Cloud Computing, Data Science & Engineering (Confluence) · 2022

In recent years, speaker recognition has become a hot research topic, especially after the DNN based embedding has started showing impressive results even for short duration audio and that too independent of the spoken text. Speaker recognition is basically used for identity recognition and with the improved results, it is now widely used in criminal investigation, voice biometric, and other interest-based customer services. In scenarios where recognition is to be done irrespective of what has been spoken, the ResNet-based architecture are able to efficiently extracts speaker embedding by implementing the residual connections to the convolutional network and standardization of the residual blocks. But their performance degrades when dealing with complex input feature space. To deal with these problems, other variants of ResNet have been experimented with. Out of these is the ResNext architecture that has been initially used for Image processing tasks and has been reported to outclass all previous Deep Neural architectures. Res2Net is known for introducing scale as another dimension in place of width and depth that improved the representation capacity of the model. This paper explores and implements Res2Net model for Speaker-recognition and verification tasks as well, on a Hindi Speakers dataset taken from Farmer’s Agri-query system and some in-house data collected from the volunteers for this task. We evaluate and compare the proposed Res2Net based systems with the ResNet model and EER has been reported which indicates the superior performance of the proposed system over the ResNet conventional model.

Read the paper · More papers on PaperTik