How Similar is Image Cognition in Humans and Computer Visions?

Hiroshi Omori, Kazunori Hanyu · 2022 Joint 12th International Conference on Soft Computing and Intelligent Systems and 23rd International Symposium on Advanced Intelligent Systems (SCIS&ISIS) · 2022

The development of deep learning for image cognition is remarkable. Computer Vision Models (CVMs) such as CNN, Vision Transformer (ViT), and CLIP, which were pretrained on a huge amount of training data, were released. We studied environmental cognition using photos. Photos were quantified by many participants manually measuring the inter-phot visual similarity (IVS). We had three types of photosets:, 200 student life photos, 242 townscapes, and 100 garden landscapes. We measured how similar the inter-photo CVM similarity is with the IVS by using MDS coordinates. In the garden landscapes, the first axis of CVM MDS was very similar to that of IVS, whereas the second axis was not similar. In Kawagoe townscapes and the student life photos, there were several clusters in IVS MDS. These clusters got spread in CVM MDS. In Kawagoe townscapes, the clusters appeared overspread. In the student life photos, both MDSs were similar, and CVM MDS looked better in fine classification. Representation seemed also involved in the differences in MDS between IVS and CVM.

Read the paper · More papers on PaperTik