Segmenting spoken language utterances into clauses for semantic classification

N.K. Gupta, Srinivas Bangalore · 2004

Robust spoken language understanding in large-scale conversational dialog applications is usually performed by classification of the user utterances into one or many semantic classes. The features used for classification are sensitive to variations caused by artifacts of spoken language, such as edits, repairs and other dysfluencies. Furthermore, the performance of these classifiers typically degrades when the user's utterance contains multiple semantic classes. In this paper, we present a semantic classification technique that first automatically removes dysfluencies and segments the user's utterance into clauses and then classifies the utterance based on the classification of the clauses. We show that this preprocessing improves the semantic classification accuracy for utterances and significantly decreases the amount of training data needed for a given classification accuracy level.

Read the paper · More papers on PaperTik