Unsupervised Speaker Identification for TV News

Daniel N. Woo, Ramazan Savas Aygün · IEEE Multimedia · 2016

Identifying the speakers in TV news would help listeners analyze and understand news content, but doing so in news videos is challenging because new faces often appear. Previous research has identified speakers on pretrained faces for TV shows and movies. Using an unsupervised method, this article proposes labeling speakers using just the available information in the news video without external information. The proposed framework segments the audio by speaker, parses closed captions for speaker names, identifies who is speaking, and performs optical character recognition for speaker names. The framework uses face recognition, face clustering, face landmarking, natural language processing tools, and speaker diarization. Results indicate 63.6 percent accuracy for identifying speakers for CNN News.

Read the paper · More papers on PaperTik