A statistical segment-based approach for spoken language understanding

Lucía Ortega, Isabel Galiano, Lluís-F. Hurtado, Emilio Sanchis, Encarna Segarra · 2010

In this paper we propose an algorithm to learn statistical language understanding models from a corpus of unaligned pairs of sentences and their corresponding semantic representation. Specifically, it allows to automatically map variablelength word segments with their corresponding semantic units and thus, the decoding of user utterances to their corresponding meanings. In this way we avoid the time consuming work of manually associate semantic labels to words, process which is needed by almost all the corpus-based approaches. We use the algorithm to learn the understanding component of a Spoken Dialog System for railway information retrieval in Spanish. Experiments show that the results obtained with the proposed method are very promising, whereas the effort employed to obtain the models is not comparable with this of manually segment the training corpus. Index Terms: Spoken Language Understanding, Semantic Classification, Statistical modelization

Read the paper · More papers on PaperTik