One approach to HTML wrappers creation

Viera Rozinajová, Ondrej Hluchý · Proceedings of the International Conference on Computer Systems and Technologies and Workshop for PhD Students in Computing · 2009

Most of the Internet sources are in unstructured HTML format, which is often difficult to process further. Content extraction from such sources is usually carried out by wrappers. There are various approaches to wrapper construction. The method proposed in this paper utilizes valued Document Object Model (DOM) trees. For each tag in the DOM tree a hash is computed from the sequence of parent tags, tag attributes, its values and the tag sub-tree. This conception ensures that same hashes will be computed for the particular wrapper even for different queries. Proposed solution has been implemented and verified.

Read the paper · More papers on PaperTik