CorpoMate: A framework for building linguistic corpora from the web

Aniruddha Adhikary, Silvia Ahmed · 2016

A linguistic corpus is a collection of an ample number of text documents serving as a data source for sampling human language, usually in computational linguistics. Conventional methods of building such a corpus involve frequent human intervention and poses difficulties during reproduction. To address the issues, the paper introduces CorpoMate, an extensible framework with a pipeline-inspired and modular architecture for automating the creation of linguistic corpora, from web resources via crawling websites or parsing feeds. It performs the necessary pre-processing as well as related tasks according to programmable queues of standard or customized tasks with easily swappable tools and, can export aggregated data into widely-accepted formats. Results from experiments performed on text processing tools and performance tests on the asynchronous, rule-based web crawling system justify the importance of having swappable tools along with the feasibility of the architecture described in the paper.

Read the paper · More papers on PaperTik