The World Wide Web as Linguistic Corpus
Charles F. Meyer, Roger J. Grabowski, Hung-Yul Han, Konstantin Mantzouranis, Stephanie Moses · 2003
Increasingly, corpus linguists have begun using the World Wide Web as a corpus for conducting linguistic analyses. The Web, however, is really a very different kind of corpus: we do not know, for instance, precisely how large it is or what kinds of texts are on it. In this chapter, we evaluate the Web as a linguistic corpus, providing estimates of its size and composition. In addition, we conduct a series of sample analyses of the Web, demonstrating that while commonly available search engines have definite limitations, they can in a matter of seconds retrieve extremely large volumes of data that are very relevant to a corpus analysis, and also provide frequency information that may not be entirely accurate but suggestive of how frequently particular words and grammatical constructions occur.