Salient Object Detection for Images Taken by People With Vision Impairments

Jarek Reynolds, Chandra Kanth Nagesh, Danna Gurari · 2024

Salient object detection is the task of producing a binary mask for an image that deciphers which pixels belong to the foreground object versus background. We introduce a new salient object detection dataset using images taken by people who are visually impaired who were seeking to better understand their surroundings, which we call VizWiz-SalientObject. Compared to seven existing datasets, VizWiz-SalientObject is the largest (i.e., 32,000 human-annotated images) and contains unique characteristics including a higher prevalence of text in the salient objects (i.e., in 68% of images) and salient objects that occupy a larger ratio of the images (i.e., on average, ∼50% coverage). We benchmarked ten modern models on our dataset. One method achieves nearly human performance while the rest struggle, mostly for images with salient objects that are large, have less complex boundaries, and lack text as well as for lower quality images. To facilitate future extensions, we share the dataset at https://vizwiz.org/tasks-anddatasets/salient-object-detection.

Read the paper · More papers on PaperTik