Using the Web as a Corpus in Natural Language Processing
Malvina Nissim, Johan Bos · 2009
Research in Natural Language Processing (NLP) has in recent years benefited from the enormous amount of raw textual data available on the World Wide Web. The presence of standard search engines has made this data accessible to computational linguists as a corpus of a size that had never existed before. Although