An output aggregation system for large scale cross-modal retrieval

Hua Yan, Jie Shao, Tian hui Hu, Zhicheng Zhao, Fei Su, Anni Cai · 2014

This paper presents our solution to MSR-Bing Image Retrieval Challenge to measure the relevance of web images and the query given in text form. We compare and integrate three typical methods (SVM-based, CCA-based, PAMIR) to conduct the large-scale cross-modal retrieval task with concept-level visual features. In SVM-based approach, the relevance of the image and the query is scored using an on-line trained SVM classifier for the query. With canonical correlation analysis (CCA), the correlations between images and queries (i.e., text) are maximized by learning a pair of linear transformations. PAMIR [1] formalizes the retrieval task as a ranking problem and introduces a learning procedure to optimize a ranking-related criterion by projecting the images to the text space. By using the concept-level visual features obtained with convolution neural network (CNN), our output aggregation system achieves 50.93% and 51.23% in terms of NDCG@ 25 on development and test data respectively.

Read the paper · More papers on PaperTik