Using Automated Core and Spurious Features Detection in Scene Recognition to Explain Computer Vision Model

Anjon Basak, Adrienne J. Raglin · 2024

The complexity of scene recognition models in intricate environments is compounded by the lack of detailed explainable techniques. Most existing methods, such as CAM and SHAP, rely on humans for detailed interpretation of the results making them unusable by downstream systems. Salient Imagenet requires human supervision and does not provide in-depth explanation with metrics in terms of human-interpretable objects. In this study, we propose an innovative approach to explain a model through automated core and spurious feature detection. Our technique quantifies the significance of objects in scenes, introducing core and spurious IoU metrics. Rather than relying on a solitary neural feature, these metrics consolidate the model’s comprehension of scene elements using TopK Neural Window Views, distinguishing between core and spurious components. Our method obviates the need for fine-tuning an object instance segmentation model like YOLO. Instead, we employ an LLM with GroundingDINO and SAM, enabling scalability to any class in the scene recognition dataset. In the experimental section, we juxtapose our approach with CAM, SHAP, and NAM, showcasing a more detailed and informative level of explainability beyond superpixel or heatmap representations. This marks a significant advancement in explaining scene recognition models, facilitating their potential utilization by downstream systems to enhance transparency.

Read the paper · More papers on PaperTik