From web page to mega-corpus: the CNN transcripts
Sebastian Hoffmann · 2007
This paper focuses on the technical and methodological issues involved in using data available on the internet as a basis for quantitative analyses of Present-day English. For this purpose, I concentrate on the creation of a specialized corpus of spoken data and outline the steps necessary to convert a large number of publicly available CNN transcripts into a format which is compatible with standard corpus tools. As an illustration of potential uses of such data, the second part of my paper then presents a sample analysis of the intensifier so. The paper concludes with a brief discussion of the advantages and limitations of this type of internet-derived data for corpus linguistic analysis.