Vision to keywords : automatic image annotation by filling the semantic gap
Junjie Zhang · UTS ePRESS (University of Technology Sydney) · 2019
its neighbours by measuring similarities among their metadata and conduct the metric learning to obtain the representations of image contents; we also generate semantic representations for images given the collective semantic information from their neighbours.In Chapter 6, the image annotation problem is addressed by routinely checking its neighbours in a graph, which is constructed by the equipped meta information of the image.We propose a graph network to model the correlations between each target image and its neighbours.To accurately capture the visual clues from the neighbourhood, a co-attention mechanism is introduced to embed the target image and its neighbours as graph nodes.Chapter 7 concludes the thesis and outlines the scope of future work.xix