A novel approach to build Kannada web Corpus

S. Parameswarappa, V. N. Narayana, G. N. Bharathi · 2012

This paper introduces the Kannada Corpus tool, a suite of Perl (Program Extraction and Reporting Language) programs implementing an iterative procedure to build Kannada corpora from the web. The procedure requires is, first a set of "seed" words list is built and later a set of “seed” URLs (Uniform Resource Locator) containing documents in the Kannada language is collected by sending queries to commercial search engines (Google and Yahoo). The obtained seeds are then used to start a crawling job using the open-source, command-line based downloading tool "wget". The downloaded documents are then processed in various ways in order to build Kannada raw corpora such as HTML (Hyper Text Markup Language) code removal, boilerplate stripping, and language identification, duplicate and near duplicate detection. We conducted an evaluation of the tool by applying it to the construction of Kannada corpora from the domains such as Recent Discussions, Articles, Recent Activities, Proverbs, Recent Feedback's, Poems and Fifteen Books, Novels, News paper, Dictionary, Blogs and Informal Chats. The results illustrate the potential usefulness of the tool.

Read the paper · More papers on PaperTik