A Real-Time Speaker Diarization System Based On Convolutional Neural Networks Architectures

Thaer M. Al-Hadithy, Mondher Frikha · 2023

Speaker diarization, which entails dividing a single audio recording into numerous voice recordings that each belong to a different speaker, is an essential step in working with audio data. Diarization systems do not recognize the speakers themselves; instead, they use unsupervised algorithms to segment audio recordings and organize utterances into speaker-specific categories. While voice-based biometric systems can recognize people, they can only be used for recordings with a single speaker. In this paper, an architecture for segment- or group-level single speaker diarization systems is proposed. The findings of the proposed algorithm's evaluation on the VoxCeleb and RAVDESS datasets, two speech databases, demonstrate that the group- level method outperforms the segment-level strategy in terms of recognition results. The difficulty of detecting numerous speakers chatting freely in an audio recording while accounting for emotional swings that may impact speech patterns is addressed in these systems, which show promise. According to the study, diarization can obtain accuracy levels of over 95.4 percent when applied to the RAVDESS and VoxCeleb datasets, which suggests that the suggested approach from speech can achieve levels of accuracy of over 96.4 percent.

Read the paper · More papers on PaperTik