Using the MGB-2 challenge data for creating a new multimodal Dataset for speaker role recognition in Arabic TV Broadcasts
Mohamed Lazhar Bellagha, Mounir Zrigui · Procedia Computer Science · 2021
Speaker role recognition is an important component in multimedia analysis for applications such as speaker naming, speaker diarization and video summarization. The lack of labeled datasets for this task has constrained algorithm evaluations. In this paper, we present a new multimodal dataset for speaker role recognition in Arabic TV programs. The dataset is artificially created using data provided by the Multi-Genre Broadcast challenge dataset. We also describe our algorithm for the processing and creation of speaker segments and their corresponding transcripts from audio documents. The spoken transcript and the speaker segments are automatically annotated for their speaker role of presenter, reporter, or a guest speaker. Based on these artificial annotations, we demonstrate for the speaker role labeling the importance of taking into account multimodal information for predicting speaker role. We present a monomodal and multimodal speaker role recognition approaches on speaker segments mined from television programs, with audio and textual classification baselines over a three-way speaker role labeling of presenter, reporter and guest.