Language Specific and Topic Focused Web Crawling

Olena Medelyan, Stefan Schulz, Jan Paetzold, Michael Poprat, Kornél Markó · 2006

The Web has been successfully explored as training and test corpus for a variety of NLP tasks ([8], [2], [6]). However, corpora derived from the Web are usually inconsistent and highly heterogeneuos in their nature, which is normally counterbalanced by extending their size to billions of words. We assume that web crawling that takes into account

Read the paper · More papers on PaperTik