DOM-based Information Extraction for the Web Sources
Xiao Li · Chinese Journal of Computers · 2002
At present, the Web becomes a major channel for people to obtain information. However, there exist inherent drawbacks in the HTML language used to represent information of Web pages. The HTML tags only tell the browsers how to display the information on the screen, but no any semantics in them. So the HTML document is not suited to be a way of data exchange for computer to process. Based on DOM and inductive learning, the paper presents a novel approach to semi automatically generate Java classes which can be dominant part of a wrapper for Web sources. The paper is an important part of the research on integrated query processing over hetergeneous data sources.