Lexical Models to Identify Unmarked Discourse Relations: Does WordNet help?
Caroline Sporleder · LDV-Forum/Journal for language technology and computational linguistics · 2008
In this paper, we address the task of automatically determining which discourse relation holds between two text spans.We focus on relations that are not explicitly signalled by a discourse marker like but.While lexical models have been found useful for the task, they are also prone to data sparseness problems, which is a big drawback given the scarcity of discourse annotated data.We therefore investigate whether the use of lexical-semantic resources, such as WordNet, can be exploited to back-off to a more general representation of lexical information in cases were data are sparse.We compare such a semantic back-off strategy to morphological generalisations over word forms, such as stemming and lemmatising.