Creating and Acquiring Electronic Texts
Susan Hockey · Oxford University Press eBooks · 2000
This chapter discusses methods for creating and acquiring electronic texts and investigates metadata requirements and ways of sampling text. The Internet is usually the starting point to look for an electronic text, normally with a general Internet search rather than with a library catalogue search. If the text is not already available in electronic form, it will either have to be keyboarded or converted into electronic form by optical character recognition (OCR). OCR systems work by making an image representation of a page of text on a scanner and then attempting to convert the image or picture of the page into text characters by recognizing each character. This chapter also surveys some corpora in common use and discusses options for the design and development of a linguistic corpus.