Challenges and opportunities in spoken document processing: Examples from keyword search and the use of prosody
Andrew Rosenberg · The Journal of the Acoustical Society of America · 2016
Spoken document processing presents challenges and opportunities when compared to text processing. Speech transcription contains errors, but speech conveys information beyond the words that are said. To deal with errors, a spoken document should be viewed as a structure of viable hypotheses not an absolute transcription. Also, the manner in which words are spoken, their prosody, can be mined for information about the speaker and his or her intent. This talk will use keyword search as an case study that requires operating under errorful and adverse conditions. It will focus on the efforts of the IBM-led LORELEI Consortium during the IARPA BABEL program over the last four and a half years. This period has been a time of rapid change in automatic speech recognition, and the BABEL program has served as a proving ground for a number of innovations including DNNs for acoustic modeling, multi-lingual acoustic features, graphemic (vs. phonemic) recognition, active learning (and reduced resource conditions), morphological analysis, and so-called “end-to-end” speech recognition. The talk will then present areas of opportunity to leverage information communicated via prosody to improve spoken document processing including emotion recognition, dialog act classification, and information extraction on spoken documents.