Fast Two-Stage Vocabulary-Independent Search In Spontaneous Speech
Peng Yu, Frank Torsten Bernd Seide · 2006
For efficient organization of speech recordings - meetings, interviews, voice mails, lectures - the ability to search for spoken keywords is an essential capability. In Seide et al. (2004) and Yu et al. (2004), we presented our work on vocabulary-independent search in spontaneous speech. That method involved linear scanning of phonetic lattices, and thus did not scale up to large collections. In this paper, we present a two-stage approach to fast search: first we retrieve segments from an index-like structure that are promising to contain the keyword, then we locate individual keyword occurrences by a detailed linear lattice scan. However, designing an efficient vocabulary-independent indexing structure is non-trivial. We use a "soft" index, similar to Allauzen et al., that provides expected term frequencies (ETF) of query terms. We propose to approximate ETF by M-gram phoneme language models estimated on the lattices (one per segment). Our index stores these language models in an inverted structure. Word spotting experiments on voicemails show that with this two-stage method, we lose under 4% FOM (figure of merit) relative at a 25-times speed-up compared with a full linear search.