Audio-visual speaker identification with multi-view distance metric learning

Haomain Zheng, Meng Wang, Zhu Li · 2010

Both audio and visual information can be useful for speaker identification in videos. This paper proposes an audio-visual speaker identification approach that benefits from a multi-view distance metric learning method. Our metric learning scheme not only builds distance measures based on the label information of training data but also the consistency of different views. In this way, better metrics can be learned in comparison with metric learning for each view individually. We conduct experiments on VidTIMIT dataset and empirical results have demonstrated the effectiveness of our approach over a set of existing methods. In addition, we also implement our method on a multi-view digit recognition task and encouraging results are also obtained.

Read the paper · More papers on PaperTik