Tagging non-native English with the TOSCA-ICLE tagger
Pieter de Haan · Corpus Linguistics and Linguistic Theory · 2000
The TOSCA-ICLE Tagging Unit (TU) has been in use for some time now to tag (part of) the material in several research centres participating in the !CLE project. The tagger has been derived from the TOSCA tagger, which was originally designed for tagging errorfree native English (as most taggers are), and was consequently trained on native material. Applying the TOSCA-ICLE tagger to the !CLE material presents us with a number of problems that can be said to be unique to learner material. This article illustrates the kinds of errors learners are apt to make, on the basis of experience with Czech, Dutch and Spanish !CLE material. These include typing errors, spelling errors, lexical and grammatical errors. A tentative classification of these errors is presented, as well as a proposed way of dealing with them.