Internet as Corpus : Automatic Construction of a Swedish News Corpus
Martin Hassel · DSpace repository (University of Tartu) · 2001
This paper describes the automatic building of a corpus of short Swedish news texts from the Internet, its application and possible future use.The corpus is aimed at research on Information Retrieval, Information Extraction, Named Entity Recognition and Multi Text Summarization.The corpus has been constructed by using an Internet agent, the so called newsAgent, downloading Swedish news text from various sources.A small part of this corpus has then been manually tagged with keywords and named entities.The newsAgent is also used as a workbench for processing the abundant flows of news texts for various users in a customized format in the application Nyhetsguiden.