Context Aware Keypoint Extraction for Robust Image Representation
Pedro Martins, P. Carvalho, Carlo Gatta · 2012
We introduce a context-aware keypoint extractor, coined as CAKE, aimed at capturing the most informative image content. We find this algorithm particularly useful in tasks such as image retrieval, scene classification, and object (class) recognition, in which local features are mainly used to provide a robust and efficient image representation. We are motivated by the fact that the majority of local feature extractors are designed to respond to a reduced number of structures. Furthermore, we observe that the existent complementarity among feature sets is often neglected. Our context-aware algorithm is designed to respond to complementary features as long as they are informative. In the particular case of images with different types of structures, one can expect a high complementarity among the features retrieved by a context-aware extractor. By contrast, images with repetitive patterns will inhibit our method from retrieving a clear summarised description of the image content. Nonetheless, the extracted set of features can be complemented with a counterpart that retrieves the repetitive elements in the image. These two cases are depicted in Figure 1. The upper image shows a context-aware keypoint extraction on a well-structured scene, which retrieves the 100 most informative keypoints. This small number of features is sufficient to provide a good coverage of the content, which includes different types of structures. The lower image illustrates the advantages of combining context-aware keypoints with strictly local ones (SFOP keypoints [2]) to obtain a better coverage of images with repetitive patterns. An information theoretic framework is used to formulate our contextaware keypoint extraction. A keypoint will correspond to a certain image location within a structure with a low probability of occurrence (high information content). For each image location x, we consider w(x) ∈RD, any viable local representation (e.g, the Hessian matrix or the structure tensor matrix) as a “codeword” that represents the neighbourhood of x. To define the saliency measure, we regard the image codewords as samples of a multivariate probability density function. We compute the probability of a codeword w(y) using a Kernel Density Estimator [4] in which the kernel is a multidimensional Gaussian function with zero mean and standard deviation σk: