Information Extraction across Linguistic Barriers

Megumi Kameyama · 1997

Information extraction (IE) systems have been tailored to extract fixed target information from documents in a fixed language. In order to be truly useful for information analysts, the target information must be user-definable and the source documents should cover multiple languages. We will map out the path toward such open-target multilingual IE systems, identifying necessary technological breakthroughs along the path. We also discuss a Japanese-English named entity extraction system under development, which represents a case of the next step along the path. Introduction: Toward Multilingual Information Extraction Systems The natural language processing field has witnessed a rapid development of the information extraction (IE) technology since the early 90's, driven by the series of Message Understanding Conferences (MUC's) in the government-sponsored TIPSTER program. 1 This technology enables a rapid, robust, and automatic extraction of certain predefined target information from ...

Read the paper · More papers on PaperTik