Putting Meaning into Your Trees

Martha Stone Palmer · Conference on Computational Natural Language Learning · 2004

The meaning of a sentence is an essential aspect of natural language understanding, yet an elusive one, since there is no accepted methodology for determining it. There is not even a consensus on criteria for distinguishing word senses. Clearly a more robust technology is needed that uses data-driven techniques. These techniques typically rely on supervised machine learning, so a critical goal is the definition of a level of semantic representation (sense tags and semantic role labels) that could be consistently annotated on a large scale. We have been training automatic WSD systems on the English sense-tagged training data based on WordNet that we supplied to SENSEVAL2 (Dang & Palmer, 2002). A pervasive problem with sense tagging is finding a sense inventory with clear criteria for sense distinctions. WordNet is often criticized for its subtle and fine-grained sense distinctions. Perhaps more consistent and coarse-grained sense distinctions would be more suitable for natural language processing applications. Grouping the highly polysemous verb senses in WordNet (on average reducing the >16 senses per verb to 8) provides an important first step a more flexible granularity for WordNet senses that improves both inter-annotator agreement (71% to 82%) and system performance (60.2% to 69%) (Dang & Palmer, 2002). The Frameset sense tags associated with the PropBank, as discussed below, provide an even more coarse-grained and easily replicable level of sense distinctions.

Read the paper · More papers on PaperTik