Using keyword spotting to help humans correct captioning faster
Yashesh Gaur, Florian Metze, Yajie Miao, Jeffrey P. Bigham · 2015
Automatic real-time captioning provides immediate and on de-mand access to spoken content in lectures or talks, and is a cru-cial accommodation for deaf and hard of hearing (DHH) people. However, in the presence of specialized content, like in techni-cal talks, automatic speech recognition (ASR) still makes mis-takes which may render the output incomprehensible. In this paper, we introduce a new approach, which allows audience or crowd workers, to quickly correct errors that they spot in ASR output. Prior approaches required the crowd worker to manu-ally “edit ” the ASR hypothesis by selecting and replacing the text, which is not suitable for real-time scenarios. Our approach is faster and allows the worker to simply type corrections for misrecognized words as soon as he or she spots them. The sys-tem then finds the most likely position for the correction in the ASR output using keyword search (KWS) and stitches the word into the ASR output. Our work demonstrates the potential of computation to incorporate human input quickly enough to be usable in real-time scenarios, and may be a better method for providing this vital accommodation to DHH people. Index Terms: speech recognition, human-computer interac-tion, spoken term detection, real-time crowd sourcing.