Transparent management of replicated WWW document clusters
Henning Pagnia, Oliver Theel, H. Schupp · 2002
It is commonly agreed that many problems of today's Internet stem from the fact that its communication links are permanently overloaded. We propose a seamlessly integratable architecture which naturally extends the currently existing World Wide Web (WWW) infra-structure by allowing users to transparently access nearby copies of replicated WWW documents. The underlying basic idea is to aggregate closely-related WWW documents into document clusters. These clusters serve as the unit for replication, meaning that all documents of a cluster are identically managed and retrieved. The benefits of this approach are reduced response times for document retrieval as well as network-wide lowered bandwidth requirements for servicing such retrieval requests. The paper describes the overall architecture of our approach and the functioning of the individual components. Additionally, details of our prototype implementation are presented together with a preliminary performance assessment based on experimental measurements.