Mid-level cues for bottom-up grouping

Tom Lee · TSpace (University of Toronto) · 2016

Bottom-up perceptual grouping is an essential but often elusive component of computer vision that occurs in support of object recognition. We base our analysis on Gestalt grouping principles that range from low-level cues, which lack contextual scope but are computationally attractive, to mid-level cues, which cover larger scope but are difficult to incorporate in a computationally efficient manner. Our thesis begins with object categorization as motivating context for perceptual grouping, highlighting the issues of complexity and learning, and the importance of feature representation in handling variability. We then make inroads into bottom-up grouping in three steps. First, we focus on the mid-level cue of symmetry, which we extend to handle curvature and taper. We make effective use of symmetry to handle a wide range of input variability, and demonstrate significant improvements from formulating symmetry-based grouping as an optimization problem. Second, we develop an energy-based superpixel grouping framework that has expressive power to accommodate multiple grouping cues that range from low-level to mid-level, and from contour-based to region-based. We demonstrate the benefit of combining multiple mid-level cues to eliminate false positives. Finally, we reformulate bottom-up perceptual grouping as a prediction task to enable us to use the framework of structured prediction to tackle grouping as a single, unified problem. We bring performance to a level competitive with recent state-of-the-art baselines to close the gap between bottom-up grouping and recognition.

Read the paper · More papers on PaperTik