The CUHK Spoken Web Search System for MediaEval 2013
Haipeng Wang, Tan Lee · 2013
This paper describes an audio keyword detection system developed at the Chinese University of Hong Kong (CUHK) for the spoken web search (SWS) task of MediaEval 2013. The system was built only on the provided unlabeled data, and each query term was represented by only one query example (from the basic set for required runs). This system was designed following the posteriorgram-based template matching framework, which used a tokenizer to convert the speech data into posteriorgrams, and then applied dynamic time warping (DTW) for keyword detection. The main features of the system are: 1) a new approach of tokenizer construction based on Gaussian component clustering (GCC) and 2) query expansion based on the technique called pitch synchronous overlap and add (PSOLA). The MTWV and ATWV of our system on the SWS2013 Evaluation set are 0.306 and 0.304. 1.