Automatic lexical acquisition based on statistical distributions

Suzanne Stevenson, Paola Merlo · 2000

We automatically classify verbs into lexical semantic classes, based on distributions of indicators of verb alternations, extracted from a very large annotated corpus. We address a problem which is particularly difficult because the verb classes, although semantically different, show similar surface syntactic behavior. Five grammatical features are sufficient to reduce error rate by more than 50% over chance: we achieve almost 70% accuracy in a task whose baseline performance is 34%, and whose expert-based upper bound we calculated at 86.5%. We conclude that corpus-driven extraction of grammatical features is a promising methodology for find-grained verb classification.

Read the paper · More papers on PaperTik