A study on out-of-vocabulary word modelling for a segment-based keyword spotting system

Alexandros Sterios Manos · DSpace@MIT (Massachusetts Institute of Technology) · 1996

The purpose of a word spotting system is to detect a certain set of keywords in continuous speech. The most common approach consists of models of the keywords augmented with "filler, " or "garbage" models, that are trained to account for non-keyword speech and background noise. Another approach is to use a large vocabulary continuous speech recognition system (LVCSR) to produce the most likely hypothesis string, and then search for the keywords in that string. The latter approach yields much higher performance, but is significantly more costly in computation and the amount of training data required. In this study, we develop a number of segment-based word spotting systems in an e ort to achieve performance comparable to the LVCSR spotter, but with only a small fraction of the vocabulary. We investigate a number of methods to model the keywords and background, ranging from a few coarse general models to refined phone representations. The task is to detect sixty-one keywords from continuous speech in the ATIS corpus. We have achieved performance of 89.8 % Figure of Merit (FOM) for the LVCSR spotter, 81.8 % using phonewords as ller models, and 79.2 % using eighteen more general models.

Read the paper · More papers on PaperTik