Multi-page list extraction: an agent-oriented approach to user-driven information extraction
Anders Arpteg · 2006
A major problem with information extraction (IE) systems is that it is difficult to handle a large number of domains. This problem can be handled in various ways. One way is to try to create a system that is as general as possible and can extraction information from a large number of domains. Another way is to create a system that can be trained to extract information. To make such a solution useful, users should not be required to spend a lot of time training or required to have expertise knowledge. If it would be possible for average users to train IE systems to extract information, then systems can be created for a large number of domains and for a large number of users. This is called a user-driven approach. A system called ISSIE has been developed to evaluate this user-driven approach, and that take advantage of semistructured information to extract relevant pieces. The system is built using the JADE agent framework, together with other tools for working with ontologies and knowledge bases. The goal with that system is to be able to train-by-example to handle a new domain. That means that a user shall be able to train the system by simply surfing the Web using a traditional browser and only give small hints about what information that are to be extracted. The system monitors the traffic and behavior of the user and then tries to automate the extraction process. By monitoring the traffic from the Web browser, and trying to understand the communication, the system is able to handle extraction from advanced Web sites. The focus in this system is to be able to extract various types of lists, e.g. a list of products at a retailer's Web site. There are currently four types of lists: singleton, simple lists, complex lists, and multipage lists. The focus in this paper is management of multipage lists. A multipage list is divided into several pages and the system must be able to navigate to and extract from these pages. A set of experiments has been conducted to evaluate the approach and the management of multipage lists. The results showed that the multipage lists extraction works well but there are some problems. However, in general, the approach is promising and shows that a user-driven approach for multi-page list extraction is viable.