Exemplar-SVMs for visual object detection, label transfer and image retrieval

Tomasz Malisiewicz, Abhinav Shrivastava, Abhinav Gupta, Alexei A. Efros · 2012

Today’s state-of-the-art visual object detection systems are based on three key components: 1) sophisticated features (to encode various visual invariances), 2) a powerful classifier (to build a discriminative object class model), and 3) lots of data (to use in largescale hard-negative mining). While conventional wisdom tends to attribute the success of such methods to the ability of the classifier to generalize across the positive class instances, here we report on empirical findings suggesting that this might not necessarily be the case. We have experimented with a very simple idea: to learn a separate classifier for each positive object instance in the dataset (see Figure 1). In this setup, no generalization across the positive instances is possible by definition, and yet, surprisingly, we did not observe any drastic drop in performance compared to the standard, category-based approaches. More specifically, we train a separate linear SVM for every exemplar in the training set (e.g., every annotated bounding box in case of object detection). Each of these Exemplar-SVMs is thus defined by a single positive instance and millions of negatives. Taken together, an Ensemble of Exemplar-SVMs (Malisiewicz et al., 2011), aims to combine the effectiveness of a discriminative object detector with the explicit correspondence offered by a nearest-neighbor approach. While each detector is quite specific to its exemplar, we empirically observe that, after a simple calibration step, an ensemble of such Exemplar-SVMs offers surprisingly good performance, roughly comparable to the much more complex latent part-based model of (Felzenszwalb et al., 2010). It is interesting to note some of the properties of

Read the paper · More papers on PaperTik