Considering HyperDocuments and Context for Indexing the Web.
Mathias Géry · 2002
The growth of the Web, with hundreds of millions users and billions of pages, gives new challenges to the Information Retrieval (IR). Most of current systems are based on a re-use of traditional models, which have been developed for textual, atomic and independents documents, and are not adapted to the Web. A promising research orientation consists in studying the impact of Web structure on indexing and querying. Some approaches use Web structure for IR, but most of them consider a \\bag-of-links", modelling the Web as a graph with HTML pages as nodes and hypertext links as edges without taking into account the links types. The HyperDocument model presented in this article is based on essential aspects of information description and comprehension: contents, composition, linear or non-linear reading and context. We present the main aspects of our Structured IR System for the Web.