Word preselection for large vocabulary speech recognition
Roberto Billi, G. Massia, F. Nesti · 2005
In this paper we describe a preselection technique, used in a large vocabulary IWR system, which imposes lexical constraints on temporal sequences of easily measured acoustic correlates, thus reducing the initial lexical uncertainty by orders of magnitude. First we give an outline of the recognition system in which this technique is used. Then the description is focused on a statistical model which accounts for the variability of the sequences of features extracted from the speech signal and which is used for lexical access. The system was evaluated on a set of 5 speakers for different choices of statistical model, training mode, kind of lexical access and vocabulary size. In particular the performance as a function of the vocabulary size was investigated. In the best condition our preselection method can reduce a 2000 word vocabulary to a set oF less than 20 words with a probability of retaining the correct word above 95%.