Perspective Chapter: What Image AI is Close to Human Senses?

Hiroshi Omori · Artificial intelligence · 2025

The development of Image AI has been remarkable. In addition to Computer Vision Models (CVMs) that use supervised learning on large-scale data such as ImageNet-21 K, CVMs that use self-supervised learning and Contrastive Language Image Pre-training (CLIP) have been developed recently. We have been researching human environmental perception using photos. We measured the visual similarity by having many participants manually classify similar photos. We had three photo sets: 100 garden landscapes, 242 cityscapes, and 200 student life photos. We investigated how closely 26 types of pre-trained CVMs matched the human sense. We also used SoftMax regression to select a synthetic CVM from these CVMs that best correlated with the visual similarity. The optimal CVM combinations varied significantly across photo sets. For garden landscapes, which only have garden photos, semantic segmentation and self-supervised CVM were effective, while for student life photos, which have a wide variety of photos, supervised CVM was effective. For cityscapes, which have intermediate variations, self-supervised CVM was effective. MDS was used to examine in detail how the optimal CVM was like human perception and how it differed. For garden landscapes, humans and the CVM agreed on the judgment of garden size but differed on the judgment of garden style. In the case of cityscapes and student life photos, the recognition of patterns in the photos was roughly consistent between humans and the CVM, but the CVM made more detailed classifications. It was also suggested that some of the differences between the two stems from human representation. Although some differences remain, we found that by combining CVMs effectively, it is possible to construct a CVM that is quite close to human senses.

Read the paper · More papers on PaperTik