Wrapping WWW information sources

Matthew Montebello · 2002

As information over the World-Wide Web (WWW) is proliferating rapidly and the demand to access it is also escalating, the need to efficiently and effectively make use of the various knowledge resources is highly imperative. Research to extract maximum benefit from the Internet is still in its infancy. Different research areas like data mining, information retrieval, machine learning and knowledge discovery, are giving the problem due attention and all seem to converge onto a common theme-wrappers. Wrappers provide access to heterogeneous information sources by converting or translating queries into source specific queries or commands. We discuss improvements on a distinctive approach, wrapper conduction, to automatically generate a wrapper once an information source has been identified. We will show how different information filtering systems we developed employ numerous Internet information sources in an attempt to exploit this knowledge base by automatically wrapping the source to safeguard the evolvability of the same system once new sources become available on the WWW. We examine and compare the performance these systems achieved when employing wrapper-conducted queries in contrast to tailored hand-coded ones; comparative results are presented.

Read the paper · More papers on PaperTik