CUHK System for the Spoken Web Search task at Mediaeval 2012
Haipeng Wang, Tan Lee · MediaEval · 2012
This paper describes our systems submitted to the spoken web search (SWS) task at MediaEval 2012. All the systems were based on a new framework which is modied from the posteriorgram-based template matching approach. This framework employs parallel tokenizers to convert audio data into posteriorgrams, and then combine the distance matrices from the posteriorgrams of dierent tokenizers to derive a combined distance matrix. Lastly dynamic time warping (DTW) is applied to the combined distance matrix to detect the possible occurrences of the query terms. For this SWS task, we used three types of tokenizers, namely Gaussian mixture model (GMM) tokenizer, acoustic segment model (ASM) tokenizer, and phoneme recognizers of rich-resource languages. Pseudo-relevance feedback (PRF) and score normalization were also used in some of the systems.