A modular open-source focused crawler for mining monolingual and bilingual corpora from the web

Vassilis Papavassiliou, Prokopis Prokopidis, Gregor Thurmair · Meeting of the Association for Computational Linguistics · 2013

This paper discusses a modular and opensource focused crawler (ILSP-FC) for the automatic acquisition of domain-specific monolingual and bilingual corpora from the Web. Besides describing the main modules integrated in the crawler (dealing with page fetching, normalization, cleaning, text classification, de-duplication and document pair detection), we evaluate several of the system functionalities in an experiment for the acquisition of pairs of parallel documents in German and Italian for the Health & Safety at work domain.

Read the paper · More papers on PaperTik