{bs,hr,sr}WaC - Web Corpora of Bosnian, Croatian and Serbian

Nikola Ljubešić, Filip Klubička · 2014

In this paper we present the construction process of top-level-domain web corpora of Bosnian, Croatian and Serbian. For constructing the corpora we use the Spi-derLing crawler with its associated tools adapted for simultaneous crawling and processing of text written in two scripts, Latin and Cyrillic. In addition to the mod-ified collection process we focus on two sources of noise in the resulting corpora: 1. they contain documents written in the other, closely related languages that can not be identified with standard language identification methods and 2. as most web corpora, they partially contain low-quality data not suitable for the specific research and application objectives. We approach both problems by using language mod-eling on the crawled data only, omitting the need for manually validated language samples for training. On the task of dis-criminating between closely related lan-guages we outperform the state-of-the-art Blacklist classifier reducing its error to a fourth. 1

Read the paper · More papers on PaperTik