Improving Visual Counterfactual Explanation Models for Image Classification via CLIP

Xiang Li, Ren Togo, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama · 2023

Deep learning models have achieved remarkable success in the field of computer vision. However, improving the quality of visual counterfactual results continues to be a significant challenge. Visual counterfactual explanation, a task that highlights image regions that need alterations to reclassify them into a different category, allows for explanations that are more intuitively understandable to humans. In this paper, we propose a method that introduces Contrastive Language-overview Image Pretraining (CLIP) as an auxiliary model to obtain a better feature pair of the query class and the distractor class, leading to more accurate visual counterfactual explanations. Experimental results on CUB-200-2011 Dataset demonstrate that our method yields a 3% improvement in Near-KP and a 0.1 increase in the "number of edits" metric when generating explanations, outperforming existing state-of-the-art methods in the image classification task.

Read the paper · More papers on PaperTik