Do CNN's features correlate with human fixations?

Marinella Iole Cadoni, Andrea Lagorio, Enrico Grosso · 2020

In recent years, CNNs are capturing the attention of a large community of researchers, attracted by the high performance of this approach and by the surprising results obtained in many recognition/classification activities. Unfortunately, the excellent performance of CNN-based systems is accompanied by a worryingly poor understanding of why they work so well. In this document, the basic mechanisms related to the extraction of points of interest (following the first convolution phases) are considered and compared to human fixations, with the aim of better understanding analogies and differences between computational models and human recognition. Alongside, the points of interest extracted from different layers are compared in order to evaluate similarity between different networks and the how interest points evolve within the same network. Human fixations and points of interest are initially used to construct density distribution maps; a novel similarity index is then proposed in order to compare these distribution maps. Experiments on the ETD database show that human fixations, on average, tend to be contained in the density of CNN points. In other words, human fixations seem to somehow optimize the active exploration of the image by properly using the information coming from the periphery of the visual field. These first interesting results could condition emerging models for visual recognition and visual indexing, pushing towards a more attentive exploitation of coarse structural information.

Read the paper · More papers on PaperTik