Towards an efficient archive of spontaneous speech: Design of computer-assisted speech transcription system
Hiroaki Nanjo, Tatsuya Kawahara · The Journal of the Acoustical Society of America · 2006
Computer-assisted speech transcription (CAST) system for making archives such as meeting minutes and lecture notes is addressed. For such a system, automatic speech recognition (ASR) of spontaneous speech is promising, but ASR results of spontaneous speech contain a lot of errors. Moreover, the ASR errors are essentially inevitable. Therefore, it is significant to design a good interface with which users can correct errors easily in order to take advantages of ASR for making speech archives. It is From these points of view that our CAST system is designed. Specifically, the system has three correction interfaces: (1) pointing device for selection from competitive candidates, (2) microphone for respeaking, and (3) keyboard. One of the most significant correction methods is selection from competitive candidates, thus more accurate competitive candidates are required. Therefore, generation methods of competitive candidates are discussed. Then, a speech recognition strategy (decoding method) based on the minimum Bayes-risk (MBR) framework is discussed. Since MBR decoding can distinguish and deal with content words and functional words, which are conventionally treated in a same manner, the approach is expected to generate more content words in competitive candidates.