Semi-Supervised Never-Ending Learning in Rhetorical Relation Identification.
Erick Galani Maziero, Graeme Hirst, Thiago Alexandre Salgueiro Pardo · Scientific Electronic Library Online (São Paulo Research Foundation, Latin American and Caribbean Center on Health Sciences Information, Conselho Nacional de Desenvolvimento Científico e Tecnológico) · 2015
Some languages do not have enough labeled data to obtain good discourse parsing, specially in the relation identification step, and the additional use of unlabeled data is a plausible solution. A workflow is presented that uses a semi-supervised learning approach. Instead of only a pre-defined additional set of unlabeled data, texts obtained from the web are continuously added. This obtains near human perfomance (0.79) in intra sentential rhetorical relation identification. An experiment for English also shows improvement using a similar workflow.