Stereo Correspondence with Local Descriptors for Object Recognition

Gee-Sern Jison · InTech eBooks · 2011

Vision2. Local region descriptors in Section 3: These descriptors transform affine invariant regions into vectors or distributions so that some distance measure can be applied to discern the similarity or difference between features.Again, those with better invariance to viewpoint changes are especially interested.3. Object modeling and recognition using local region descriptors from multi-view in Section 4: A couple methods are reviewed that develop models by combining the information from local descriptors extracted across multiple views.These methods offer good examples on how to integrate local invariant features across different views.4. A case study on performance evaluation and benchmark databases in Section 5: Implementation of others' methods for performance comparison with one's own proposed method takes a tremendous amount of time and efforts.Therefore, a database commonly accepted for performance benchmark is needed, and different methods can be evaluated on the same testbed.A performance evaluation example is reviewed with an introduction on its database, followed by a snapshot on other databases also good for study on 3D object recognition using stereo correspondences. Affine regions for stereo correspondenceAffine-invariant region detectors can identify the affine-invariant regions on multiple images which are the projections of the same 3D surface patches.The regions are also considered as covariant with geometric and photometric transformations, as the regions detected in one image can be mapped onto those detected in the other using these transformations.Different affine detectors give different local regions in terms of different locations, sizes, orientations and the numbers of detected regions.Mikolajczyk et al. (2005) have evaluated six affine region detectors, including Harris-affine, Hessian-affine, edge-based region, intensity extrema-based region, salient region and maximally stable extremal region (MSER).This evaluation focuses on the performance of matching between two images with variations caused by viewpoint, scale, illumination, blur and JPEG compression.The detectors for regions only covariant to similarity transform are excluded in their evaluation, for example the interest regions extracted to develop the Scale-Invariant Feature Transform (SIFT) by Lowe (1999;2004) and the scale invariant features by Mikolajczyk & Schmid (2001).However, the SIFT descriptor (Lowe, 1999;2004) is used in this evaluation to characterize the intensity patterns of the regions detected by the above six detectors.The scope of this chapter is on finding stereo correspondences for object recognition, subject to the requirement that the object's model is built on at least a pair of stereo images with different viewpoints.In certain cases, the objects in stereo or multiple images may appear slightly different in scale.Therefore the detectors that perform better than others in rendering correct matches under viewpoint and scale changes are of special interest in this chapter.This performance can be justified by the repeatability and matching score from the evaluation in Mikolajczyk et al. (2005).It is shown that the Harris-affine detector, Hessian-affine detector and the maximally stable extremal region (MSER) detector are three promising ones in offering reliable stereo correspondences under viewpoint and scale changes.Note that illumination changes, blur and JPEG compression are among the major challenging parameters when recognizing a test image, the three aforementioned detectors also perform well when testing against these parameters, as revealed by Mikolajczyk et al. (2005).

Read the paper · More papers on PaperTik