Topic Segmentation

Matthew Purver · 2011

This chapter discusses the task of topic segmentation: automatically dividing single long recordings or transcripts into shorter, topically coherent segments. First, it looks at the task itself, the applications which require it, and some ways to evaluate accuracy. The chapter explains the most influential approaches - generative and discriminative, supervised and unsupervised - and discusses their application in particular domains. It also focuses on the techniques for understanding on a fine-grained, bottom-up level: identifying fundamental units of meaning or interactional structure, such as sentences, named entities and dialogue acts. Benchmark datasets is only a useful task when applied to recordings of some length - short segments of speech such as an utterance in a typical spoken dialogue system tend already to be topically homogeneous and thus not to require segmentation. As a result, it only started to receive attention once long recordings became available. Controlled Vocabulary Terms audio recording; interactive systems

Read the paper · More papers on PaperTik