Supervised Learning of Lexical Semantic Verb Classes Using Frequency Distributions
Suzanne Stevenson, Paola Merlo, Natalia Kariaeva Rutgers · 1999
We report a number of computational experiments in supervised learning whose goal is to automatically classify a set of verbs into lexical semantic classes, based on frequency distribution approximations of grammatical features extracted from a very large annotated corpus. Distributions of five syntactic features that approximate transitivity alternations and thematic role assignments are sufficient to reduce error rate by 56% over chance. We conclude that corpus data is a usable repository of verb class information, and that corpusdriven extraction of grammatical features is a promising methodology for automatic lexical acquisition. 1 Introduction Recent years have witnessed a shift in grammar development methodology, from crafting large grammars, to annotation of corpora. Correspondingly, there has been a change from developing rule-based parsers to developing statistical methods for inducing grammatical knowledge from annotated corpus data. The shift has mostly occu...