Investigation of speaker embeddings for cross-show speaker diarization

Michael Rouvier, Benoît Favre · 2016

This paper proposes to investigate speaker embeddings, a representation extracted from hidden layers of deep neural networks trained on a speaker identification task, on cross-show diarization. The new representation brings an improvement over i-vectors, and we show that while shallow hidden layers give best results on the single-show condition, deeper layers yield better performance on cross-show diarization. This confirms that deep representations model higher level features which help generalizing to different acoustic conditions. Experiments, conducted on the French corpus of REPERE, show that the deep speaker embeddings technique decreases DER by 0.82 points.

Read the paper · More papers on PaperTik