Information Extraction from Multi-Document Threads

David J. Masterson · 2003

Information extraction (IE) is the task of extract-ing fragments of important information from natu-ral language documents. Most IE research involves algorithms for learning to exploit regularities inher-ent in the textual information and language use, and such systems generally assume that each document can be processed in isolation. We are extending IE techniques to multi-document extraction tasks, in which the information to be extracted is distributed across several documents. For example, many kinds of work-flow transactions are realized as sequences of electronic mail messages comprising a conversation among several participants. We show that IE perfor-mance can be improved by harnessing the structural and temporal relationships between documents. 1

Read the paper · More papers on PaperTik