A study of generic models for unsupervised on-line speaker indexing

Soonil Kwon, Shrikanth Shri Narayanan · 2004

On-line speaker indexing sequentially detects the points where a speaker identity changes in a multi-speaker audio stream, and classifies each speaker segment. The paper addresses two challenges. The first relates to monitoring, which requires on-line processing. The second relates to the fact that the number/identity of the speakers is unknown. The indexing needs to be made in an unsupervised process. To address these issues, we apply a predetermined generic speaker-independent model set, sample speaker model (SSM). This set can be useful for more accurate speaker modeling and clustering without requiring training models on target speaker data. Once a speaker-independent model is selected from the sample models, it is adapted into a speaker-dependent model progressively. Experiments were performed with the speaker recognition benchmark NIST Speech (1999). Results showed that our new technique, simulated using the Markov chain Monte Carlo method, gave 92.47% indexing accuracy on telephone conversation data.

Read the paper · More papers on PaperTik