Hybrid Network with Multi-Level Global-Local Statistics Pooling for Robust Text-Independent Speaker Recognition
Woo Hyun Kang, Jahangir Alam, Abderrahim Fathan · 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) · 2021
In this paper, we propose a new hybrid system for extracting a speaker embedding vector. More specifically, the proposed system employs a multi-level global-local statistics pooling method in order to aggregate the speaker information within short time-span and utterance-level context. In order to evaluate the proposed system, a set of experiments on the NIST SRE 2016, Short-duration speaker verification (SdSV) Challenge 2021, and VoxCeleb datasets were conducted, and the proposed hybrid network was able to outperform the conventional approaches trained on the same dataset. Moreover, our experiments showed that the proposed system is able to achieve stable performance even when using a relatively smaller dataset, which highlights the efficiency of the proposed system in extracting the speaker-dependent information.