Grammar fragment acquisition using syntactic and semantic clustering

Kazuhiro Arai, Jeremy H. Wright, Giuseppe Riccardi, Allen L. Gorin · 1998

A new method for automatically acquiring Fragments for understanding fluent speech is proposed. The goal of this method is to generate a collection of Fragments, each representing a set of syntactically and semantically similar phrases. First, phrases observed frequently in the training set are selected as candidates. Each candidate phrase has three associated probability distributions: of following contexts, of preceding contexts, and of associated semantic actions. The similarity between candidate phrases is measured by applying the Kullback-Leibler distance to these three probability distributions. Candidate phrases that are close in all three distances are clustered into a Fragment. Salient sequences of these Fragments are then automatically acquired, and exploited by a spoken language understanding module to classify calls in AT&T's "How May I Help You?" task. These Fragments allow us to generalize unobserved phrases. For instance, they detected 246 phrases in the test-set that we...

Read the paper · More papers on PaperTik