Merging search spaces for subword spoken term detection
Timo Mertens, Daniel M. Schneider, Joachim Köhler · 2009
Abstract We describe how complementary search spaces, addressed bytwo different methods used in Spoken Term Detection (STD),can be merged for German subword STD. We propose fuzzy-search techniques on lattices to narrow the gap between sub-word and word retrieval. The first technique is based on anedit-distance, where no a priori knowledge about confusions isemployed. Additionally, we propose a weighting method whichexplicitly models pronunciation variation on a subword leveland thus improves robustness against false positives. Recall isimproved by 6% absolute when retrieving on the merged searchspace rather than using an exact lattice search. By modelingsubword pronunciation variation, we increase recall in a high-precision setting by 2% absolute compared to the edit-distancemethod.Index Terms: subword speech recognition, spoken term detec-tion, pronunciation variation 1. Introduction As vast amounts of media can be stored digitally, methods arerequired to make them searchable. For data containing speech,a challenging task is Spoken Term Detection (STD) in whichall occurrences of a word or phrase have to be located preciselyin a database.A straightforward approach to STD is to use a large vocabularycontinuous speech recognizer (LVCSR) to transcribe thespeech parts of the data. However, this implies that if a wordis not in the vocabulary it cannot be transcribed, and thuscannot be found upon search time. This phenomenon is knownas the out-of-vocabulary (OOV) problem and is especiallyprominent in languages that make use of compounding oremploy complex morphological patterns to create new words.Examples of such languages are Finnish, German and Turkish[1, 2, 3]. One way to alleviate this problem is to use subwordsas the recognition unit. In contrast to words, the inventory ofsubwords, e.g. syllables, is finite and hence known a priori.Their compositional nature allows for retrieval of compoundterms that otherwise could not be found on the word level.However, a decrease in recall compared to word STD has beenobserved [2, 4].The 1-best transcription given by the LVCSR will containother misrecognitions besides those caused by OOVs. This isespecially the case for spontaneous speech, which is often un-planned and may contain a high number of false starts. SeveralSTD techniques have emerged to cope with these recognitionerrors. One approach is retrieval on word or subword lattices(or lattice-like variants) instead of the 1-best transcription [5].In [4] we compared this approach to an edit-distance searchon the 1-best transcription [6], which matches the query withthe 1-best transcription and calculates a distance according toa similarity measure. Like other fuzzy approaches to subwordSTD [7], this method can tolerate a degree of mismatchbetween the recognition output and the query, and is thus ableto find keywords in erroneous transcriptions. On a GermanSTD task, we found that on a subword level, each approachcovers a different area of the search space, and comes withrelative advantages and disadvantages.In this contribution we present ways of merging the searchspaces covered by the two previously mentioned retrievalapproaches. The aim is to exploit the complementary nature ofsubword lattice and subword fuzzy-search by combining theminto a single, unified search method. Further, we investigatehow explicit modeling of pronunciation variation on a subwordlevel can improve robustness of the retrieval method.