WebCorp: an integrated system for web text search

Antoinette Renouf, Andrew Kehoe, Jayeeta Banerjee · 2007

The web has unique potential to yield large-volume data on up-to-date language use, obvious shortcomings notwithstanding. Since 1998, we have been developing a tool, WebCorp, to allow corpus linguists to retrieve raw and analysed linguistic output from the web. Based on internal trials and user feedback gleaned from our site (http://www. webcorp.org.uk/), we have established a working system which supports thousands of regular users world-wide. Many of the problems associated with the nature of web text have been accommodated, but problems remain, some due to the non-implementation of standards on the Internet, and others to reliance on commercial search engines, which mediation slows up average WebCorp response time and places constraints on linguistic search. To improve WebCorp performance, we are in the process of creating a tailored search engine, an infrastructure in which WebCorp will play an integral and enhanced role. In this paper, we shall give a brief description of WebCorp, the nature and level of its current functionality, the linguistic and procedural problems in web text search which remain; and the benefits of replacing the commercial search engine with tailored websearch architecture.

Read the paper · More papers on PaperTik