Semantic Data Mining of Short Utterances

Lee Begeja, Harris Drucker, David C. Gibbon, Patrick Haffner, Zhu Liu, Bernard Renger, Behzad Shahraray · 2006

This paper introduces a methodology for speech data mining along with the tools that the methodology requires. We show how they increase the productivity of the analyst who seeks relationships among the contents of multiple utterances and ultimately must link some newly discovered context into testable hypotheses about new information. While in its simplest form, one can extend text data mining to speech data mining by using text tools on the output of a speech recognizer, we have found that it is not optimal. We show how data mining techniques that are typically applied to text should be modified to enable an analyst to do effective semantic data mining on a large collection of short speech utterances.

Read the paper · More papers on PaperTik