Combining hierarchical clustering and machine learning to predict high-level discourse structure
Caroline Sporleder, Alex Lascarides · 2004
We propose a novel method to predict the interparagraph discourse structure of text, i.e. to infer which paragraphs are related to each other and form larger segments on a higher level.Our method combines a clustering algorithm with a model of segment "relatedness" acquired in a machine learning step.The model integrates information from a variety of sources, such as word co-occurrence, lexical chains, cue phrases, punctuation, and tense.Our method outperforms an approach that relies on word co-occurrence alone.