Top-down Gamma Saliency - Learning to Search for Objects in Complex Scenes
Ryan Burt, José Carlos Príncipe · 2018
Saliency measures are often used to predict fixation location in images. However, a pure bottom up saliency is not useful for visual search in a complex scene with many objects since it is only driven by the input image. Alternatively, neural networks can localize objects within scenes, but rely on a brute force classification of heuristic bounding boxes. We propose a top-down attention mechanism that combines the traditional saliency measures with the learned ability of neural networks to distinguish between objects. To do this, we will use a set of feature maps produced by the convolutional layers of a trained classification network as the inputs to our saliency measure instead of a traditional RBG or LAB image. On top of these feature maps, we can learn a set of weights to bias the saliency towards specific objects. We test this top-down approach against the traditional bottom-up approach in a synthetic environment where it proves to be more adept at finding specific objects quickly in crowded scenes.