Estimating confidence using word lattices
Thomas Kemp, Thomas K. Schaaf · 1997
For many practical applications of speech recognition systems, it is desirable to have an estimate of confidence for each hypothesized word, i.e. to have an estimate which words of the speech recognizer's output are likely to be correct and which are not reliable. Many of today's speech recognition systems use word lattices as a compact representation of a set of alternative hypothesis. We exploit the use of such word lattices as information sources for the measure-of-confidence tagger JANKA [1]. In experiments on spontaneous human -to-human speech data the use of word lattice related information significantly improves the tagging accuracy. 1. INTRODUCTION Current speech recognition systems are far from perfect. Unfortunately, number and location of the errors in their output is usually unknown. However, this information could be used in a number of applications. Examples are word selection for unsupervised adaptation schemes like MLLR [5], automatic weighting of additional, non-speec...