Reducing the effect of OOV query words by using morph-based spoken document retrieval

Ville T. Turunen · 2008

Morph-based spoken document retrieval uses morpheme-like subword units for both language modeling and as index terms. Problems of out-of-vocabulary (OOV) words are avoided as the morph recognizer can recognize any word in speech as a sequence of subwords. The effect of previously unseen query words (i.e. words that are not in the language model training text) is analyzed for Finnish spoken document retrieval. The performance of the morph-based system is compared to a wordbased approach. Language models with artificially high OOV query word rates are built and the results show that morphbased retrieval suffers significantly less from the OOV query words than word-based. Extracting alternative recognition candidates from confusion networks further improves the results, especially for morph-based retrieval.

Read the paper · More papers on PaperTik