Crowdsourcing Thumbnail Captions via Time-Constrained Methods
Carlos Aguirre, Amama Mahmood, Chien‐Ming Huang · 2022
Speech interfaces, such as personal assistants and screen readers, employ captions to allow users to consume images; however, there is typically only one caption available per image, which may not be adequate for all settings (e.g., browsing large quantities of images). Longer captions require more time to consume, whereas shorter captions may hinder a user’s ability to fully understand the image’s content. We explore how to effectively collect both thumbnail captions—succinct image descriptions meant to be consumed quickly—and comprehensive captions, which allow individuals to understand visual content in greater detail. We consider text-based and time-constrained methods to collect descriptions at these two levels of detail, and find that a time-constrained method is most effective for collecting thumbnail captions while preserving caption accuracy. We evaluate our collected captions along three human-rated axes—correctness, fluency, and level of detail—and discuss the potential for model-based metrics to perform automatic evaluation.