Concordancing the web: promise and problems, tools and techniques

William H. Fletcher · 2007

The web is an inexhaustible reservoir of machine-readable texts in most of the world’s written languages for compiling corpora or consulting directly as a ‘corpus’. This paper first surveys some characteristics of the web and discusses the potential rewards and practical limitations of exploiting the web either directly as a linguistic corpus or to compile corpora. Particular attention is paid to search engines, our gateways to the web. The author then reviews several innovative applications of web data to corpus-related issues. KWiCFinder (KF), developed by the author to help realize the web’s promise for language scholars and learners, is described and motivated in detail. KF, readily accessible to novices yet powerful enough for advanced researchers, conducts web searches, retrieves matching online documents, and produces an interactive keyword in context concordance of the search terms. This paper then discusses the pitfalls of ‘webidence’ in serious research and proposes an initial solution. Finally the author reviews the future of the web for corpus research and application.

Read the paper · More papers on PaperTik