Domain-specific Web site identification: the CROSSMARC focused Web crawler

Konstantinos Stamatakis, Vangelis Karkaletsis, Γεώργιος Παλιούρας, James Horlock, Claire Grover, James Curran, Shipra Dingare · Edinburgh Research Explorer (University of Edinburgh) · 2003

This paper presents techniques for identifying domain specific Web sites that have been implemented as part of the EC-funded R&D project, CROSSMARC. The project aims to develop technology for extracting interesting information from domain-specific Web pages. It is therefore important for CROSSMARC to identify Web sites in which interesting domain specific pages reside (focused Web crawling). This is the role of the CROSSMARC Web crawler. 1.

Read the paper · More papers on PaperTik