Statistical resolution of scope ambiguity in natural language
Galen Andrew, Bill MacCartney · 2004
A crucial obstacle is that there is no labeled corpus available that is suitable for learning to resolve scope ambiguities. As a consequence, we’ve had to generate our own data set by (a) selecting sentences appropriate to our purposes, and (b) hand-labeling the sentences selected. We recognize that this represents a methodological compromise: our results would carry more weight if the training data, and especially the test data, had not been generated by the researchers themselves. However, we’ve attempted to mitigate this shortcoming by establishing in advance clear guidelines for selecting and labeling sentences. We obtained our data from sentences drawn from GRE and LSAT logic games. We chose this limited domain because these sentences are specifically designed to have an interpretation that is unambiguous to a human reader. Thus there is less subjectivity about the correct labelings, and determining the correct scoping is less likely to depend on context and pragmatics.