Verb sense and verb subcategorization probabilities
Douglas Roland, Daniel S. Jurafsky · Natural language processing · 2002
this paper we measure these probabilities only for syntactic argument frames, but the Lemma Argument Probability hypothesis bears equally on the semantic/thematic expectations shown by studies such as Ferreira and Clifton (1986) and Trueswell et al. (1994). Our results also suggest that the subcategorization frequencies that are observed in a corpus result from the probabilistic combination of the lemma's expectations and the probabilistic effects of context. The other important implication of these two sources of variation is methodological. Our results suggest that, because of the inherent differences between isolated sentence production and connected discourse, probabilities from one genre should not be used to normalize experiments from the other. In other words, `test-tube' sentences are not the same as `wild' sentences. We also show that seemingly innocuous methodological devices, such as beginning Verb Sense and Verb Subcategorization Probabilities 3 sentences-to-be-completed with proper nouns (Debbie remembered...) can have a strong effect on resulting probabilities. Finally, we show that such frequency norms need to be based on the lemma or semantics, and not merely on shared orthographic form. 2 Methodology