An Effective Web Crawling Algorithm by Graph Search Techniques
김진일, 권유진, Sung‐Ryeol Kim, 박근수 · 한국정보과학회 학술발표논문집 · 2008
Web crawlers ar fundamental software which iteratively downloads web pages by following links of web pages starting from a small set of initial URLs. Previously several web crawling orderings have been studied by [1,2]. In this paper we consider various graph search techniques and treat internal links and external links in different manners for effective crawling. For maximum cardinality search (MCS) and lexicographic breadth-first search (LexBFS), we present linear time algorithms based on the partition refinement method [3,4,5]. The experimental results show that maximum cardinality search is preferable to other graph search techniques for web crawlers since it visits web pages with high PageRank earlier than others.