Transcribing continuous speech using mismatched crowdsourcing
Preethi Jyothi, Mark Hasegawa‐Johnson · 2015
Mismatched crowdsourcing was recently proposed as a poten-tial approach to deriving moderately accurate speech transcrip-tions using crowd workers unfamiliar with the language be-ing spoken. In introducing this approach, we demonstrated its promise with the help of an isolated word recovery task for Hindi. However, it remained open whether mismatched crowd sourcing can yield non-trivial accuracy in a continuous speech task. In this work, we focus on this question and demonstrate a word error rate of under 45 % in a large-vocabulary task (again for Hindi). In achieving this, we develop several new tech-niques capable of scaling effectively to continuous speech. We also provide an information theoretic analysis and estimate the amount of information lost in transcription by the mismatched crowd workers to be under 5 bits.