PARC 3.0: A Corpus of Attribution Relations

Silvia Pareti · 2016

Quotation and opinion extraction, discourse and factuality have all partly addressed the annotation and identification of Attribution Relations.However, disjoint efforts have provided a partial and partly inaccurate picture of attribution and generated small or incomplete resources, thus limiting the applicability of machine learning approaches.This paper presents PARC 3.0, a large corpus fully annotated with attribution relations (ARs).The annotation scheme was tested with an inter-annotator agreement study showing satisfactory results for the identification of ARs and high agreement on the selection of the text spans corresponding to its constitutive elements: source, cue and content.The corpus, which comprises around 20k ARs, was used to investigate the range of structures that can express attribution.The results show a complex and varied relation of which the literature has addressed only a portion.PARC 3.0 is available for research use and can be used in a range of different studies to analyse attribution and validate assumptions as well as to develop supervised attribution extraction models.

Read the paper · More papers on PaperTik