UCBN: A New Audio-Visual Broadcast News Corpus for Multimodal Speaker Verification Studies

Girija Chetty, Michael Wagner · University of Canberra Research Portal · 2006

The performance of face, voice, and multimodal speaker verification systems in complex and non-controlled scenarios, is typically lower than systems developed in highly controlled environments. With the aim to facilitate the development of robust multi-modal speaker recognition systems, a new multi-modal (audio-visual) Australian broadcast UCBN (University of Canberra Broadcast News) corpus was developed by capturing about 30 hours of television daily news program from several free-to-air Australian TV channels over a period of two years. In this paper we describe the acquisition of UCBN, and a new video preprocessing technique used for detection of newscasters and anchor person shots in news sequences. The speaker verification experiments using feature fusion of acoustic and visual speech features extracted from the mouth region are also reported. The performance of the complex UCBN database is compared with that of the controlled VidTIMIT database.

Read the paper · More papers on PaperTik