Exploring Self-Supervised Representations for Text-Dependent Speaker Verification
Sankala Sreekanth · 2024
This paper presents team speech information processing lab at IITH (SIPLAB-IITH) submission to the text-dependent speaker verification challenge 2024. Deep embedding-based methods are widely used in text-independent speaker verification (TI-SV). However, they are less explored in textdependent speaker verification (TD-SV) because labeled data for TD-SV is limited. Given the success of self-supervised models in downstream tasks with minimal labeled data, we investigate the use of self-supervised representations for deep embedding-based TD-SV systems. Additionally, we investigate the approach of combining separate TI-SV systems with password verification models, both using self-supervised representations, for the TD-SV task. First, the password verification model rejects trials with incorrect passwords. The remaining trials are then scored using the TI-SV system. Finally, our submission to the challenge achieved a minDCF of 0.029 in Track 1 (ranked $1^{\text {st }}$) and 0.142 in Track 2 (ranked $3^{\text {rd }}$).