On the Use of Cross- and Self-Module Attentive Statistics Pooling Techniques for Text-Independent Speaker Verification

Jahangir Alam · 2023

In neural speaker verification, statistics pooling plays a key role in the learning and extraction of a speaker embedding vector. In this contribution, we perform an investigative study on the use of cross-module and self-module attention statistics pooling mechanisms to extract speaker discriminant utterance level embeddings. More specifically, we propose a novel hybrid neural network that employs a2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) and a TDNN-Long Short Term Memory (LSTM) hybrid network in a parallel manner. In order to capture the complementarity between two parallelly connected modules, the proposed system also make use of a cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context. We conduct a set of experiments on the Voxceleb and CNCeleb corpora, and the proposed approach is able to provide better performance than the conventional approaches trained on the same dataset.

Read the paper · More papers on PaperTik