Theory and applications in information extraction from unstructured text
Tianhao Wu · 2002
Information Extraction (IE) is a rapidly growing field in natural language processing in part because there is much data available through the Internet, and much of the data, such as free text and semi-structured text, needs to be preprocessed before computers can automatically understand them. The focus of this thesis is in the area of IE systems. I first review background and prior work in IE. I compare IE with full text understanding and with Information Retrieval (IR). Furthermore, I introduce several basic learning methods and the evaluation metrics used in IE systems. I also briefly describe IE related work resulting from the Message Understanding Conferences [MUCs] as well as the HDDJTM Collection Builder. Following this, I present two IE systems constructed as part of this research. I then review transformation-based error-driven learning, a technique that is a powerful learning method in the field of IE. I have used transformation-based erro!driven learning to discover patterns for extraction of thread conversation starts in chat data. This outperforms a human expert approach based on the commonly employed F p metric.