Learning Visual Classifiers using Human-centric Annotations.

Ishan Misra, C. Lawrence Zitnick, Margaret A. Mitchell, Ross Girshick · arXiv (Cornell University) · 2015

When human annotators are given a choice about what to label in an image, they apply their own subjective judgments on what to ignore and what to mention. We refer to these noisy human-centric annotations as exhibiting human reporting bias. Examples of such annotations include tags and keywords found on photo sharing sites, or in datasets containing captions. In this paper, we use these noisy annotations for learning visually correct classifiers. Such annotations do not use consistent vocabulary, and miss a significant amount of the information present in an image; however, we demonstrate that the noise in these annotations exhibits structure and can be modeled. We propose an algorithm to decouple the human reporting bias from the correct visually grounded labels. Our results are highly interpretable for reporting what's in the image versus what's worth saying. We demonstrate the algorithm's efficacy along a variety of metrics and datasets, including MS COCO and Yahoo Flickr 100M. We show significant improvements over traditional algorithms for both classification and captioning, doubling the performance of existing methods in some cases.

Read the paper · More papers on PaperTik