Modelling Salient Object-Object Interactions to Generate Textual Descriptions for Natural Images

Hossein Adeli, Babak Yadranjiaghdam, Nathan Pool, Nasseh Tabrizi · 2016

In this paper, we propose a new method for automatically generating textual descriptions of images. Our method consists of two main steps: Using saliency maps, it detects the areas of interests in the image, and then creates the description by recognizing the interactions between detected objects within those areas. These interactions are modeled using the pose (body parts configuration) of the objects as a reference. To create sentences, a syntactic model is used that builds sub-trees around the detected objects and then combines those sub-trees using recognized interaction. Our results show the improved accuracy and naturalness of the generated descriptions.

Read the paper · More papers on PaperTik