Creating a Web Corpus Using GO
M. Kucic · 2021
The Web contains large amounts of textual data which could be used as a source to create new corpora, yet there are not many plug and play solutions for scraping specific parts of the websites. This paper presents a new open-source solution for downloading and parsing HTML websites which can be configured from one configuration file. As a demonstration of this method, a new ad hoc corpus was built. The corpus contains a total of 2,395,735 titles and articles collected from 14 most popular Croatian websites.