Extracting Time and Space Relations from Natural Language Text

Jennifer D’Souza · Zenodo (CERN European Organization for Nuclear Research) · 2015

Relation extraction is a core task in natural language processing that concerns the extraction of relations among the entities and events mentioned in a text document. Despite the vast amount of work on relation extraction from text, there has been relatively little work that focuses on understanding how entities and events are temporally and spatially related. This dissertation examines two key tasks in relation extraction, temporal relation extraction and spatial relation extraction. Temporal relation extraction involves determining the temporal ordering over events, dates, and other temporal entities. We focus on fine-grained temporal relation extraction, where we classify a pair of temporal entities as belonging to one of a predefined set of up to 14 temporal relation types. We propose a knowledge-rich, hybrid approach to this task. Specifically, we employ sophisticated linguistic knowledge derived from a variety of semantic and discourse relations, and leverage a hybrid system combining the strengths of rule-based and learning-based approaches. Experiments on newswire and medical data show that our approach yields a relative error reduction of about 15% over the state of the art. Spatial relation extraction, on the other hand, concerns the determination of how spatial entities are related to each other. While previous work on this task has focused on extracting relations involving stationary spatial entities, we examine a more challenging version of the task, in which we additionally identify spatial relations involving objects in motion. Unlike in many relation extraction tasks where exactly two entities can participate in a relation, in spatial relation extraction involving objects in motion, up to eight spatial entities can participate. To handle the complexity of extracting these relations, we propose a multi-pass sieve approach to spatial relation extraction, which can capture the partial dependencies among spatial entities without sacrificing computational tractability. When evaluated on a newly released corpus, our approach significantly outperforms state-of-the-art spatial relation extraction systems.

Read the paper · More papers on PaperTik