A New Unsupervised Short-Utterance based Speaker Identification Approach with Parametric t-SNE Dimensionality Reduction
Omar Elnaggar, Roselina Arelhi · 2019
State-of-the-art speaker identification (SI) systems have achieved accuracies of 100% with long-duration utterances which are impractical. Recently, short-utterance based systems have gained attention although identification rates are lower. This paper presents an approach for text-dependent speaker phonemebased (<; 1sec) SI with parametric t-distributed stochastic neighbor embedding (pt-SNE) for dimensionality reduction of features to provide 3D-visualization. The approach employs Gaussian mixture model enhanced by K-means++ and gap statistic methods. As there is no other similar work, a fair comparison is unavailable. The 75% rate achieved is comparable to other works using (i) short-utterances (ii) pt-SNE for recognition of other data types.