SiamTDNN: Enhancing Discriminative Embeddings for Speaker Diarization

Runqing Zhang, Huijun Lu, Dunbo Cai, Zhiguo Huang, Yujian Du, Ling Qian, Yijun Zhang · Journal of Circuits Systems and Computers · 2023

Recent advances in speaker embeddings promote a great development of speaker diarization. However, determining ‘who spoke when’ in the meeting scenarios is still challenging due to similar speaker voices and unknown speaker quantity. In this paper, this research proposes enhanced discriminative features for speaker diarization, including discriminative speaker-specific features based on Siamese networks, and a speaker re-verification method. With Siamese architecture, SiamTDNN, this research first explores latent representations which is capable of modeling intra-class and inter-class differences between speakers, by training with audio pairs. Then, the re-verification method is introduced with a local-global strategy to identify speakers in a multi-person talking scene. Our method provides a novel speaker embedding with enhanced discriminative power for disambiguated speakers and achieves an elevated upper bound on the number of speakers. The proposed speaker embedding achieved an EER of 1.1% and a minDCF of 0.1192 on VoxCeleb1 for the speaker verification task. Extensive experiments on AiShell-4, ICSI, AMI and VoxConverse demonstrate the effectiveness of the proposed method with an average DER reduction of 3% and an RTF of 0.0792.

Read the paper · More papers on PaperTik