Spoken document retrieval based on phoneme recognition

Martin Wechsler · 1998

Recently, vast amounts of audio and video material containing spoken information have become available in digital format, for example radio news recordings. Along with this development there is an increased demand to retrieve spoken information in response to a user's information need. This thesis addresses the problem of spoken document retrieval (SDR). A particular goal is to perform experimental studies on documents spoken in German. The approach taken by this thesis requires a phoneme recognizer, which initially generates phoneme sequences from the spoken documents. The main issues of phoneme-recognition--based SDR are (1) missing word boundaries in the phoneme sequences and (2) the high number of phoneme recognition errors, requiring an error-tolerant method to detect query words. We present Probabilistic String Matching (PSM); a new retrieval method where query words are spotted in document phoneme sequences that are corrupted by recognition errors. This method includes the detec...

Read the paper · More papers on PaperTik