Comparing part-based representations of deep convolutional neural networks with those of human vision through representational similarity analysis
Jiaqi Adam Huang, Peter C. Gerhardstein · 2021
Multiple theories of human object recognition argue for the importance of semantic parts in the formation of intermediate representations. However, the role of semantic parts in Deep Convolutional Neural Networks (DCNN), which encapsulate the most recent and successful computer vision models, is poorly examined. We extract representations of DCNNs corresponding to differential performance with stimuli in which different parts of the same exemplar are deleted, and then compare these representations with those of human observers obtained in a behavioral experiment, using representational similarity analysis (RSA). We find that DCNN representations correlate strongly with those of observers, while acknowledging that these DCNN representations may not be part-based given an equally high correlation between DCNN output and part size. Additionally, the exemplars incorrectly identified by DCNNs tend to have less “human-like” representations, which demonstrates RSA as a potential novel method for interpreting error in intermediate processes of recognition of DCNNs.