Improving the identification of non-anaphoricitusing support vector machines
Jose Carlos Clemente Litran, Kenji Satou, Kentaro Torisawa · 2004
Identification of non-anaphoric use of the pronoun it is crucial to achieve full anaphora resolution. Nevertheless, this problem has been either ignored or considered too simple to deserve a deeper study. In this paper we present a machine-learning approach using Support Vector Machines. We collected several instances of both anaphoric and non-anaphoric it from the GENIA corpus, together with syntactic information about the context. We show how by using a limited amount of knowledge our approach can achieve better accuracy than previous methods. We also analyze the relevance of features used to predict non-anaphoric uses.