Research and Implementation of Distributed and Multi-topic Web Crawler System
Jinlin Wang · Jisuanji gongcheng · 2009
This paper proposes an architecture of distributed Web crawler system based on data-trapper.It implements a multi-topic schema based on classics-label,so that one crawler can contain different topics adaptively and designs a two-tiered weighted task partition algorithm that realizes target-guided URL configuration based on Agents’ load while providing better dynamic scalability.It improves URL storage with Trie tree,which efficiently supports URL search,insertion and repetition judgment.