Bayesian analysis of similarity matrices for speaker diarization

Alexey Sholokhov, Timur Pekhovsky, Oleg Kudashev, Шулипа Андрей Константинович, Tomi Kinnunen · 2014

Inspired by recent success of speaker clustering in Total Variability space we propose a new probabilistic model for speaker diarization based on Bayesian modeling of pairwise similarity scores. The recordings are represented by symmetric similarity matrices of likelihood ratio scores from probabilistic linear discriminant analysis (PLDA) trained on short-term i-vectors. We employ Bayesian approach to address the problem of unknown number of speakers in conversation. Diarization error rates on the NIST 2008 SRE telephone data indicate comparable performance with state-of-the-art eigenvoice-based diarization. But unlike the eigenvoice approach, our method finds the number of speakers automatically, making the proposed model more viable for practical applications.

Read the paper · More papers on PaperTik