Identification of speakers in movie dialogs using audiovisual cues
Ying Li, Shrikanth Shri Narayanan, C.‐C. Jay Kuo · IEEE International Conference on Acoustics Speech and Signal Processing · 2002
The problem of identifying speakers from a movie dialog scene is addressed in this paper. While most previous work on speaker identification has been carried out using pure audio data, more robust results could be obtained by integrating the knowledge from multiple media sources such as visual and audio information when they are available. In this work, we first identify and isolate speech segments from background by applying video shot detection, audio classification and adaptive silence detection techniques, then a decision is made based on the calculated likelihood between the incoming speech data and pre-trained speaker/background models. Moreover, to verify the effectiveness of the adaptive silence detector, we have compared it with a statistically trained silence model. Experimental results show that the proposed algorithm can achieve approximately 84% identification accuracy by integrating multiple media cues.