The accidental corpus: some issues in extracting linguistic information from the Web

Antoinette Renouf, Andrew Kehoe, David Mezquiriz · 2004

The Web is a text store which can potentially supplement traditional corpora as a source of up-to-date linguistic data. The WebCorp project investigates this potential, and in its second year tackles some residual problems inherent in the nature of Web text, thereby refining its retrieval and analysis tool for the facilitation of corpus linguistic study.

Read the paper · More papers on PaperTik