Speaker Recognition System based on Identity Vector using T-SNE Visualization and Mean-shift Algorithm
Kourosh Kiani, Atefeh Baniasadi · 2019
The process of manually labeling data is not affordable. Moreover, the lack of labeled data has led to a big performance gap between scoring baseline techniques in speaker recognition. This paper aims to propose two separate systems to fill this gap. The first system uses the t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm to represent the unlabeled development i-vectors into two space dimensions, then cluster them using the mean-shift algorithm. Finally, the Within-Class Covariance Normalization (WCCN) algorithm and test normalization technique are applied to remove unwanted variability from i-vectors. In the second system, zero normalization is also employed on the baseline system released by the NIST 2014 i-vector challenge dataset. The cosine similarity has been computed for scoring in the proposed methods. The evaluation results on the NIST 2014 i-vector challenge dataset show that the proposed methods achieve 23% and 8% relative improvement of the minimum detection cost function (minDCF) respectively. Moreover, we obtained 25% improvement by fusing these systems.