Causal Inference for Confounder-Purify Vision Transformers
Dengfeng Yue, Jia Zou, Xiaoxiao Jin, Tuo Leng · 2024
Obtaining genuine causal relationships, rather than spurious correlations, is a critical challenge in computer vision. Vision Transformers (ViT) and its variants have gained popularity in the field due to their exceptional ability to capture global context correlations through the self-attention mechanism. However, the inherent risk of capturing spurious correlations influenced by confounders poses a concern. In response, we propose Confounder-purify Vision Transformers based on causal inference and generative adversarial networks (GANs), which is able to capture genuine causal relationships. Our training process incorporates a Confounder-purify (CP) module, ensuring that ViT and its variants acquire causal features unaffected by confounders. We evaluate our approach on datasets such as FairFace, UTKFace, and NICO, and demonstrate the contrasting changes in attention maps before and after the removal of confounding factors using Grad-CAM. The results show that our method effectively reduces the impact of confounders on results and improves the accuracy of the model.