On the influence of regularization techniques on label noise robustness: Self-supervised speaker verification as a use case
Abderrahim Fathan, Xiaolin Zhu, Jahangir Alam · 2024
Clustering-based Pseudo-Labels (PLs) are widely used to optimize Speaker Embedding networks and train Self-Supervised Speaker Verification (SV) systems. However, this self-supervised training scheme relies on highly accurate PLs. In this paper, we perform a large investigative study of the effect of several regularization techniques (mixup, label smoothing, employing sub-centers) on the label noise robustness of self-supervised speaker verification systems. We study these techniques and apply them to various recent metric learning loss functions for better generalization of self-supervised speaker verification systems. In particular, we investigate the effect of these losses and regularizations on the robustness of the self-supervised SV task against label noise using various clustering models to generate real-world PLs of different noise patterns and levels. We provide a thorough comparative analysis of the generalization performance of these losses and regularization techniques using different numbers of clusters and propose some combination systems that are effective against label noise and lead to considerable improvements in SV performance.