Advancing Speaker Diarization With Whisper Speech Recognition for Different Learning Environments

Aarsh Desai, N.V.J.K Kartik, Priyesh Gupta, Vinayak Vinayak, T. S. Ashwin, Manjunath K Vanahalli, Ramkumar Rajendran · 2024

This study investigates the application of speaker diarization methodologies within educational contexts, addressing the unique linguistic and environmental challenges present in Indian learning environments. Despite the frequent switching between English and local languages, current state-of-the-art models often fail to accurately process these multilingual interactions. To bridge this gap, we propose integrating the diarization outputs of Pyannote.audio with the transcription capabilities of OpenAI’s Whisper model, alongside a custom Voice Activity Detection (VAD) and embedding clustering pipeline. Using a dataset of 137 recordings from an online Human-Computer Interaction (HCI) course and an augmented reality (AR) classroom, our findings highlight the enhanced performance of the proposed methodologies. The Pyannote-Whisper pipeline achieved the lowest Diarization Error Rate (DER) of 0.26, outperforming other models, including commercial applications such as Deepgram and Otter. The custom VAD + Embedding Clustering Pipeline also demonstrated robustness in multilingual contexts, showcasing its potential for accurate diarization without relying on transcription. Fine-tuning and hyperparameter optimization were performed, yet did not significantly improve upon the proposed models, indicating the robustness of the initial approach. This study underscores the potential for adapting existing models to improve speaker diarization in multilingual educational settings, providing a significant step towards more robust analysis of collaborative learning interactions.

Read the paper · More papers on PaperTik