Speaker Clustering Based on Minimum Rand Index

Wei-Ho Tsai, Hsin‐Min Wang · 2007

This paper presents an effective method for clustering unknown speech utterances based on their associated speakers. The proposed method jointly optimizes the generated clusters and the number of clusters by estimating and minimizing the Rand index of the clustering. The Rand index, which reflects clustering errors that utterances from the same speaker are placed in different clusters, or utterances from different speakers are placed in the same cluster, reaches its minimal value only when the number of clusters is equal to the true speaker population size. We approximate the Rand index by a function of the similarity measures between utterances and employ the genetic algorithm to determine the cluster where each utterance should be located, such that the overall clustering errors are minimized. The experimental results show that the proposed speaker-clustering method outperforms the conventional method based on hierarchical agglomerative clustering in conjunction with the Bayesian information criterion to determine the number of clusters.

Read the paper · More papers on PaperTik