Crowdsourcing for Speech Transcription
Gabriel Parent · 2013
This chapter is written for people with little or no previous experience with crowdsourcing to guide them through the design, creation, and publishing of tasks for labeling and transcribing speech utterances. Because no two readers will have exactly the same requirements for their annotation, this chapter addresses topics that apply to a large range of tasks. Based on the literature and on personal experience, some guidelines are provided: how to achieve good-quality transcriptions, how to interact with the workers, how much to pay, and so on. The reader is guided through the steps that lead to crowdsourced transcriptions: audio preprocessing, task design, submitting the open call and postprocessing for quality control. Using audio CAPTCHA to transcribe speech is a good example of crowdsourcing. Recognition Output Voting Error Reduction (ROVER) is the most widely used form of aggregation for crowdsourced transcriptions. Crowdsourcing has several advantages over conventional transcription.