Scalable collection summarization and selection
Ron A. Dolin, Divyakant Agrawal, E. El Abbadi · 1999
Information retrieval o ver the Internet increasingly requires the ltering of thousands of i n f o r mation sources.As the number and variety of sources increases, new ways of automatically summarizing, discovering, and selecting sources relevant t o a u s er's query a re needed.Pharos i s a h i g hly scalable distributed architecture for locating heterogeneous information sources.Its design is hierarchical, thus allowing it to scale well as the numberofinformation sources increases.We demonstrate the feasibility of the Pharos architecture using 2500 Usenet newsgroups as separate collections.Each n e wsgroup is summarized viaautomated Library o f C o ngress classi cation.We s how that using Pharos as an i n termediate retrieval mechanism provides acceptable accuracy of source selection compared to selecting sources using complete classi cation information, while maintaining good scalability.This implies that hierarchical distributed metadata and automated classi cation are potentially useful paradigms to address scalability problems in large-scale distributed information r etrieval applications.