Instance Sampling Methods for Pronoun Resolution

Holger Wunsch, Sandra Kübler, Rachael Cantrell · Recent Advances in Natural Language Processing · 2009

Instance sampling is a method to balance extremely skewed training sets as they occur, for example, in machine learning settings for anaphora resolution. Here, the number of negative samples (i.e. non-anaphoric pairs) is usually substantially larger than the number of positive samples. This causes classifiers to be biased towards negative classification, leading to suboptimal performance. In this paper, we explore how dierent techniques of instance sampling influence the performance of an anaphora resolution system for German given dierent classifiers. All sampling methods prove to increase the F-score for all classifiers, but the most successful method is random sampling. In the best setting, the F-score improves from 0.541 to 0.608 for memory-based learning, from 0.561 to 0.611 for decision tree learning and from 0.511 to 0.584 for maximum entropy learning.

Read the paper · More papers on PaperTik