Word-lattice based spoken-document indexing with standard text indexers

Frank Torsten Bernd Seide, Kit Thambiratnam, Roger Peng Yu · 2008

Indexing the spoken content of audio recordings requires automatic speech recognition, which is as of today not reliable. Unlike indexing text, we cannot reliably know from a speech recognizer whether a word is present at a given point in the audio; we can only obtain a probability for it. Correct use of these probabilities significantly improves spoken-document search accuracy.

Read the paper · More papers on PaperTik