Current Issues in Markup-Based Knowledge Extraction
Udo Kruschwitz · 2002
Extracting content from Web pages can be useful for a number of reasons. Our motivation is to help a user in the search for documents in subdomains of the Web such as company sites and intranets. Unlike online product catalogues, the data sources we are interested in are of heterogenous nature. A model that reflects the underlying semantic structure of the document collec-tion can be very helpful. However, it is difficult to get hold of a domain model that can easily be plugged into such a system. We have been working on this prob-lem for some time now and this paper will report our ongoing work in the field of markup-based knowledge extraction. Markup is used to identify conceptual in-formation. This enables us to build a simple domain