Deep Structured Models for Text Understanding

Yifan Su · Repository for Publications and Research Data (ETH Zurich) · 2016

A key problem for unstructured text understanding is to identify the relevant actors in each document and their semantic relationship.Jointly solving various text semantic understanding tasks has recently raised researchers' attention.While structured graphic models encountered a fair amount of success in NLP in the past decades, they have recently been superseded by deep neural networks.This thesis is concerned with developing tools to jointly solve multiple entity analysis tasks using deep structured model.Particularly, we focus on two core NLP tasks that have shown to be interdependent: first, coreference resolution, which consists of clustering together mentions of the same entity within a document, and second, entity disambiguation (ED), which involves resolving entities in a source document to a target entry in a knowledge base such as Wikipedia.Our independent systems can efficiently learn neural semantic features with deep learning for both tasks, which performs quite well in ED but not as well in coreference.One of the challenges in coreference resolution is that it requires taking into consideration many linguistic phenomena: constraints arise from syntax, semantics, discourse, and pragmatics.The scenario makes it difficult to build effective learningbased resolution systems without using hand-crafted features.We then show that a set of simple features inspecting surface lexical properties together with our neural semantic features based on vector embedding can capture a range of these effects and make efficient, high-performing independent baseline systems for both coreference and entity disambiguation.To exploit world knowledge of semantics, we then turn to the task of ED and tackle it jointly with coreference.On the one hand, the coreference model can learn knowledge from a broader resource base (e.g.Wikipedia).Our ED model, on the other hand, can draw on information from multiple mentions of the same coreference cluster we are attempting to resolve.All parameters in both individual models can be automatically backpropagated and optimized on a self-defined structured loss function through the entire network, making it possible to learn more informative feature representation.At prediction time, we employ approximate inference to find the best resolution sequence for all mentions within a document.We see performance gains over each independent model, which suggests that deeper integration of various parts of the NLP stack is a natural way to yield a better understanding of the text.

Read the paper · More papers on PaperTik