Generating text description from content-based annotated image

Yan Zhu, Hui Yu Xiang, Wenjuan Feng · 2012

This paper proposes a statistical generative model to generate sentences from an annotated picture. The images are segmented into regions (using Graph-based algorithms) and then features are computed over each of these regions. Given a training set of images with annotations, we parse the image to get position information. We use SVM to get the probabilities of combinations between labels and prepositions, obtain the data to text set. We use a standard semantic representation to express the image message. Finally generate sentence from the xml report. In view of landscape pictures, this paper implemented experiments on the dataset we collected and annotated, obtained ideal results.

Read the paper · More papers on PaperTik