Adaptation of Visual Models with Cross-modal Regularization - eScholarship

Costa Pereira, Caterina María · 2015

Semantic representations of images have been widely adopted in ComputerVision. A vocabulary of concepts of interest is first identified and classifiers arelearned for the detection of those concepts. Images are classified and mapped to aspace where each feature is a score for the detection of a concept. This representationbrings several advantages. First, the generalization from low-level featuresto concept-level enables similarity measures that correlate much better with userexpectations. Second, because semantic features are, by definition, discriminantfor tasks like image categorization, the semantic representation enables a solutionfor such tasks with low-dimensional classifiers. Third, the semantic representationis naturally aligned with recent interest on contextual modeling. This is ofimportance for tasks such as object recognition, where detection of contextuallyrelated objects has been shown to improve detection of certain objects of interest,or semantic segmentation, where the coherence of segment semantics can be exploitedto achieve more robust segmentations. Lastly, due to their abstract nature,semantic spaces enable a unified representation for data from different contentmodalities, e.g. images, text, or audio. This opens up a new set of possibilities formultimedia processing, enabling operations such as cross-modal retrieval, or imagede-noising by text regularization. This unified representation for multi-modal datais the starting point of the proposed framework on adaptation of visual modelswith cross-modal regularization.We start by pointing the problems in computing similarity on heterogeneousdata, proposing two fundamental hypotheses to deal with those issues. One,learning a space that maximizes the correlation on the (heterogeneous) data; two,learning a representation where data lies at a higher level of abstraction. Empiricalevidence is shown in favor of each hypothesis; furthermore the hypotheses areshown to be complementary. We follow on the (semantic) abstraction hypothesisfor a deeper understanding on the robustness of these representations and to studythe richness of this space, as it highly influences the discriminative power of suchdescriptors.It has been shown that categories unknown to the semantic space, whenrepresented in it, exhibit a pattern of co-occurring concepts that describe them accuratelyand sensibly; e.g. the concept of fishing might not belong to the semanticspace and instead be represented by the set water, boat, people and gear. Eventhough the amount of labeled data continues to increase with ongoing efforts fromdifferent research communities, it is a challenging task to build a semantic spacethat is universal. We show evidence towards robustness of representations in thesemantic space.Noting that images are frequently published on the web together withloosely related text, we use the semantic representations described above to introducethe theoretical principles to a feature regularizer for image semantic representationsbased on auxiliary data. This proves very effective on improving retrievalprecision and recall in the task of content-based image retrieval (CBIR). It’s resultsare compared to recently developed methods, achieving significant gains inthree benchmark datasets, raising the bar of state-of-the-art performance for imageretrieval.

Read the paper · More papers on PaperTik