UWCL at MediaEval 2013: Similar Segments in Social Speech Task
Gina‐Anne Levow · 2013
This paper describes the participation of the University of Washington Computational Linguistics Laboratory (UWCL) in the Similar Segments in Social Speech task at MediaEval 2013. Participants in this task develop systems that, given a span of speech from a recorded conversation, aim to identify all and only highly similar regions in other recordings. As this was a new task for this year, the goal was to establish a baseline and a framework for future experimentation. The approach aimed to address two particular challenges posed by the task: the lack of prior segmentation of the conversations and the limited material provided by a single brief example segment. To this end, the system employed a query-by-example information retrieval framework using passage retrieval to identify segments dynamically and query expansion to support robust retrieval. Query expansion provided substantial gains when applied to both manual and automatic transcriptions; results using automatic transcripts were competitive with those using manual ones.