ResNet Does Not Perform Intelligent Understanding of Picture in Image Recognition
Allegra Allgeier, Austin Au-Yeung, Heinrich Matzinger · 2022 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA) · 2022
We present failure cases of ResNet-50 on car damage detection, damage location classification, and bear image classification. When trained on images from the Internet, we find that 1) close-ups and cropped photos of undamaged cars are classified as damaged, 2) damage location classification may depend more on the photo angle than the actual damage location, 3) images with bears cropped out are classified as containing bears, and 4) bear images cut into squares and randomly reshuffled are still classified as containing bears. We deduce from these experimental results and a theoretical consideration on convolutions that ResNet does not take into account global relative positions of patterns. In other words, there is no recognition of a global shape. Instead, ResNet relies on spurious correlations between certain features and the associated classes. For example, in our car damage recognition experiment, we see that ResNet never truly recognizes a car damage or its location but bases its prediction on features such as how close to the car and from which angle the photo was taken. We explain that, depending on the task, removal of such spurious correlations can be nearly impossible. From our experience, the methods presented here are common in the industry. Therefore, we want to caution practitioners who believe ResNet and possibly other CNN architectures are intelligently understanding an image when in fact they are not.