Development of Web Crawler to Build Indonesian Text Corpus

Janson Hendryli, Viny Christanti Mawardi · IOP Conference Series Materials Science and Engineering · 2020

Abstract Recent improvement in natural language understanding research can be attributed to the availability of large scale datasets. Those datasets are mainly in English. In this work, we develop a web crawler with the purpose of extracting Indonesian news content from the DetikNews website and building a large dataset of texts. The web crawler is developed by following the waterfall model using Python, Scrapy, and BeautifulSoup4. It collects more than 790k news from DetikNews, spanning from 2011 to 2020, which consists of a total number of more than 190 million words, almost 2 million unique words, and more than 14 million sentences.

Read the paper · More papers on PaperTik