On the Use of Cross- and Self-Module Attentive Statistics Pooling Techniques for Text-Independent Speaker Verification
Jahangir Alam · 2023
In neural speaker verification, statistics pooling plays a key role in the learning and extraction of a speaker embedding vector. In this contribution, we perform an investigative study on the use of cross-module and self-module attention statistics pooling mechanisms to extract speaker discriminant utterance level embeddings. More specifically, we propose a novel hybrid neural network that employs a2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) and a TDNN-Long Short Term Memory (LSTM) hybrid network in a parallel manner. In order to capture the complementarity between two parallelly connected modules, the proposed system also make use of a cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context. We conduct a set of experiments on the Voxceleb and CNCeleb corpora, and the proposed approach is able to provide better performance than the conventional approaches trained on the same dataset.