CNN Image Recognition is Mainly Based on Local Features

Heinrich Matzinger, Allegra Allgeier · 2022

We show that image classification with Convolutional Neural Networks (CNNs) follows a different strategy than how humans recognize objects: while a human may consider the global shape or a 3D picture of the object, a CNN’s decision is based mainly on recognizing local-patterns without taking into account their relative positions. We show this by feeding a CNN an image that has been cut into many squares and randomly reassembled. We observe that the original object is still recognized in the scrambled image with high probability up to a certain point. This allows us to estimate the size of the patterns used for recognizing objects by the CNN. We also show that for this analysis to make sense, we need to introduce an alternative class, "other" which contains a very vast amount of different motives, too vast for the CNN to learn.

Read the paper · More papers on PaperTik