"Sheldon speaking, Bonjour!"
Hervé Bredin, Anindya Roy, Nicolas Pécheux, Alexandre Allauzen · 2014
We address the problem of speaker identification in multimedia data, and TV series in particular. While speaker identification is traditionally a supervised machine-learning task, our first contribution is to significantly reduce the need for costly preliminary manual annotations through the use of automatically aligned (and potentially noisy) fan-generated transcripts and subtitles.