CEA LIST@imageCLEF 2013: Scalable Concept Image Annotation

Hervé Le Borgne, Adrian Stefan Popescu, Amel Znaidia · CLEF (Working Notes) · 2013

We report the participation of the CEA LIST to the Scalable Concept Image Annotation Subtask of ImageCLEF 2013. The full system is based on both textual and visual similarity to each concept, that are merged by late fusion. Each image is visually represented with a bag of visterm, computed from a dense grid of SIFT every 3 pixels, that a locally soft coded and max pooled on a codebook of size 1024 and spatially extended with a pyramid 1 1 + 3 1 + 2 2, resulting into a vector of size 8192. The visual neighbors of a query are given by the L1 distance to the images of the learning database. The similarity of a query to one of the 95/116 concepts to identify is the sum of the similarity of each neighbor to the concept. The decision is set at 1 for all concepts above one standard deviation from the average similarity to all concepts for the considered query. The similarity between a training image and a concept is computed from an intermediate vectorial representation of its tags. We tested several spaces of representation for the tags, including wikipedia concepts sorted according to their popularity or characterized according FlickR data. As well, the size of space was pruned at several values, from 5000 to 200; 000. The tag representation are max-pooled to make the training image vec- tor, such that the resulting vector contain the maximal similarity to each concept of the intermediate space. The 96/116 concepts to identify are represented in the same space and their similarity to the training image is the cosine between the intermediate space representation. As well, we ranked all training images to each concept to identify in order to learn visual classiers (linear SVM). We tested several strategies to

Read the paper · More papers on PaperTik