Using Visual Information to Predict Lexical Preference
Shane Bergsma, Randy G. Goebel · 2012
Most NLP systems make predictions based solely on linguistic (textual or spoken) input. We show how to use visual information to make better linguistic predictions. We focus on selectional preference; specifically, determining the plausible noun arguments for particular verb predicates. For each argument noun, we extract visual features from corresponding images on the web. For each verb predicate, we train a classifier to select the visual features that are indicative of its preferred arguments. We show that for certain verbs, using visual information can significantly improve performance over a baseline. For the successful cases, visual information is useful even in the presence of cooccurrence information derived from webscale text. We assess a variety of training configurations, which vary over classes of visual features, methods of image acquisition, and numbers of images. 1