Exploiting Debate Portals for Semi-Supervised Argumentation Mining in User-Generated Web Discourse
Ivan Habernal, Iryna Gurevych · 2015
Analyzing arguments in user-generated Web discourse has recently gained atten-tion in argumentation mining, an evolving field of NLP. Current approaches, which employ fully-supervised machine learn-ing, are usually domain dependent and suffer from the lack of large and diverse annotated corpora. However, annotating arguments in discourse is costly, error-prone, and highly context-dependent. We asked whether leveraging unlabeled data in a semi-supervised manner can boost the performance of argument component identification and to which extent is the approach independent of domain and reg-ister. We propose novel features that ex-ploit clustering of unlabeled data from de-bate portals based on a word embeddings representation. Using these features, we significantly outperform several baselines in the cross-validation, cross-domain, and cross-register evaluation scenarios. 1