Words and Word Usage: Newspaper Text versus the Web
Vinci Liu, James Curran · 2005
This paper explores the differences in words and word usage in two corpora – one derived from newspaper text and the other from the web. A corpus of web pages is compiled from a controlled traversal of the web, producing a topicdiverse collection of 2 billion words of web text1. We compare this Web Corpus with the Gigaword Corpus, a 2 billion word collection of news articles. The Web Corpus is applied to the task of automatic thesaurus extraction, obtaining similar overall results to using the Gigaword. The quality of synonyms extracted for each target word is dependent on the word’s usage in the corpus. With many more words available on the web, a much larger Web Corpus can be created to obtain better results in different nlp tasks.