Information retrieval and extraction from the web: the CROSSMARC approach
Vangelis Karkaletsis, Constantine D. Spyropoulos · 2004
The paper presents the CROSSMARC approach for the complex task of identification of interesting web sites and web pages and the extraction of information from them. This task is hard because most of the information on the Web today is in the form of HTML documents, which are designed for presentation purposes and not for automatic extraction systems. This task becomes even harder in a multilingual context, where web pages in different languages need to be considered. CROSSMARC approach focuses on the easy customization of web information retrieval and extraction technology to new domains and languages. This is achieved by adopting and implementing an open, multi-lingual and multi-agent architecture that integrates the CROSSMARC components into a web-based prototype system, as well as by providing an infrastructure that facilitates customization of its components to new domains and languages.