Extending Self-Distilled Self-Supervised Learning For Semi-Supervised Speaker Verification

Jeong-Hwan Choi, Jehyun Kyung, Ju-Seok Seong, Ye-Rin Jeoung, Joon‐Hyuk Chang · 2023

In this study, we extend self-distillation with no labels (DINO), a successful self-supervised learning framework, by combining it with supervised classification (SC) for semi-supervised speaker verification with limited labeled data. We introduce a transfer learning framework that pre-trains and fine-tunes the encoder using DINO and SC, respectively, and a multitask learning framework that shares the encoder while having separate projection layers for both methods. To achieve lower inter-speaker similarity, we propose a joint learning framework sharing both the encoder and projection layer for DINO and SC. We also propose an auxiliary contrastive loss between embeddings derived from labeled and unlabeled utterances and introduce a two-stage learning strategy to apply margin penalty effectively. Experimental results on the VoxCeleb corpus indicate that the joint learning framework outperforms the other frameworks and is closest to achieving the performance of fully supervised learning.

Read the paper · More papers on PaperTik