Collecting Speech Data using Amazon's Mechanical Turk for Evaluating Voice Search System

李 清宰, Tatsuya Kawahara, Rudnicky Alexander · 2011

This paper describes a crowd-sourcing method to collect speech data using Amazon’s Mechanical Turk (MTurk). We designed a task (HIT) to collect speech data as an evaluation set for voice search and another task to verify the quality of the collected speech data. More than a thousand utterances are collected very efficiently. It turned out that more than 90% of them are valid with correct transcript, and reasonable recognition accuracy is achieved. Using the data, we conducted evaluation of the voice book search system, and confirmed that the combination of slot-based vector space models provides higher search accuracy than the conventional single vector space model.

Read the paper · More papers on PaperTik