Rethinking explainable AI: The gap between saliency-based explanation and user understanding for object detection models

Ruoxi Qi, Guoyang Liu, Jindi Zhang, Janet Hui-wen Hsiao · International Journal of Human-Computer Studies · 2026

Saliency-based explainable AI (XAI) methods are commonly used to explain the behaviors of AI models, despite the limited research on whether such methods can indeed enhance user understanding. Here we proposed a set of tasks to systematically and objectively evaluate user’s global understanding of object detection models at the feature, object, and image levels. We found that while presenting AI’s hits, misses, and false alarms to users could enhance feature-level and some aspects of object-level understanding, presenting saliency-based explanations could not provide any additional help and did not help direct user’s attention to relevant features. Meanwhile, presenting AI’s hits, misses, and false alarms alone did not help users distinguish AI’s hits from misses and did not enhance image-level understanding. At the image level, among the participants, assuming that AI would behave like themselves appeared to be the best strategy for predicting AI’s behavior, since any attempts to revise such assumption resulted in further deviations from AI’s actual behaviors. Thus, it is necessary to develop more effective XAI methods, particularly for object detection models. Our eye movement analyses showed that participants who used similar strategies to AI also tended to perform more similarly to AI, suggesting that we could instruct users to use their own strategy as a reference point to predict AI’s behavior accordingly. Also, participants’ eye movement consistency and attention strategy similarity to AI’s were associated with different aspects of user understanding, suggesting that eye movements could be used as non-intrusive measures to monitor user understanding for providing user-specific explanations in future XAI methods.

Read the paper · More papers on PaperTik