Voice Processing for Police Hotlines
Bassel Nasr, Maroun Chamoun, Jean Marc Steyaert · 2023
This paper explores the potential of applying machine learning techniques to enhance the capabilities of emergency services operation. The study examines the preprocessing techniques employed and delves into the application of speaker diarization for identifying multiple speakers in audio input, comparing the performance of different models in this task. The authors develop a custom speaker diarization model tailored to specific scenarios, capable of effectively handling voice recordings with significant speaker changes. Additionally, the study investigates various machine learning models for speech recognition, ultimately selecting XLSR- Wav2vec2, a model published by Facebook, for further analysis. The chosen model exhibits high accuracy in identifying spoken words and facilitates the automation of dataset creation in both Arabize (a combination of Latin script and Arabic numerals) and Arabic languages. The findings of this study contribute to the advancement of machine learning in call center applications, particularly in speech analysis and diarization tasks. The presented results shed light on the feasibility and performance of machine learning techniques in the hotlines response domain, offering valuable insights for researchers and practitioners.