Learning Categorical Shape from Captioned Images
Tom S.H. Lee, Sanja Fidler, Alex Levinshtein, Sven Dickinson · 2012
Given a set of captioned images of cluttered scenes containing various objects in different positions and scales, we learn named contour models of object categories without relying on bounding box annotation. We extend a recent language-vision integration framework that finds spatial configurations of image features that co-occur with words in image captions. By substituting appearance features with local contour features, object categories are recognized by a contour model that grows along the object's boundary. Experiments on ETHZ are presented to show that 1) the extended framework is better able to learn named visual categories whose within-class variation is better captured by a shape model than an appearance model, and 2) typical object recognition methods fail when manually annotated bounding boxes are unavailable.