Information Structure Prediction for Visual-world Referring Expressions
Micha Elsner, Hannah Rohde, Alasdair D. F. Clarke · 2014
We investigate the order of mention for objects in relational descriptions in visual scenes.Existing work in the visual domain focuses on content selection for text generation and relies primarily on templates to generate surface realizations from underlying content choices.In contrast, we seek to clarify the influence of visual perception on the linguistic form (as opposed to the content) of descriptions, modeling the variation in and constraints on the surface orderings in a description.We find previously-unknown effects of the visual characteristics of objects; specifically, when a relational description involves a visually salient object, that object is more likely to be mentioned first.We conduct a detailed analysis of these patterns using logistic regression, and also train and evaluate a classifier.Our methods yield significant improvement in classification accuracy over a naive baseline.