Automatically Detecting and Attributing Indirect Quotations

Silvia Pareti, Tim O’Keefe, Ioannis Konstas, James Curran, Irena Koprinska · 2013

Direct quotations are used for opinion mining and information extraction as they have an easy to extract span and they can be attributed to a speaker with high accuracy.However, simply focusing on direct quotations ignores around half of all reported speech, which is in the form of indirect or mixed speech.This work presents the first large-scale experiments in indirect and mixed quotation extraction and attribution.We propose two methods of extracting all quote types from news articles and evaluate them on two large annotated corpora, one of which is a contribution of this work.We further show that direct quotation attribution methods can be successfully applied to indirect and mixed quotation attribution.* *These authors contributed equally to this work.by quotation marks, which makes them easy to extract.However, annotated resources suggest that direct quotations represent only a limited portion of all quotations, i.e., around 30% in the Penn Attribution Relation Corpus (PARC), which covers Wall Street Journal articles, and 52% in the Sydney Morning Herald Corpus (SMHC), with the remainder being indirect (Ex.1c) or mixed (Ex.1b)quotations.Retrieving only direct quotations can miss key content that can change the interpretation of the quotation (Ex.1b) and will entirely miss indirect quotations.

Read the paper · More papers on PaperTik