Standardizing wordforms in a spoken corpus
Gerald Nelson · Literary and Linguistic Computing · 1997
In this paper, the problems of standardizing the wordforms in an orthographically transcribed spoken corpus are described. In the first part, the status of wordforms in orthographic transcriptions and the problem of variant wordforms are described, and the consequences of not introducing some measure of standardization into the spoken transcriptions are pointed out. In the second part, the procedures we used for standardizing the wordforms in the British ICE corpus (ICE-GB) are described. Finally, the ways in which variant wordforms may be dealt with during text retrieval are discussed.