Automated Text Analysis: Cautionary Tales
Catherine N. Ball · Literary and Linguistic Computing · 1994
The increasing availability of electronic text and text analysis tools has made it possible to analyse vast amounts of data in a short amount of time. However, natural language processing is not a solved problem, and even large research systems representing decades of development do not perform at the level of human language processors. Since such systems are not sufficiently robust for general use, most literary and linguistic corpus analysts make use of heuristics and simple tools for text analysis. But while such ‘shallow’ approaches offer improvements in speed and accuracy over traditional manual methods, there are many pitfalls for the unwary. In this paper we consider some pitfalls and temptations that attend the automated analysis of large text corpora: sample size, the recall problem, analysing only what is easy to find, and counting what is easiest to count. We suggest that, given the state of the art in text processing tools, such tools must be used with a full awareness of their limitations, and should be coupled with or replaced by manual methods when appropriate.