Explore fine-grained discriminative visual explanation when making classification decision

Jianyi Wan, Zhengxia Gao, Aiwen Jiang · 2018

Language and image are two most important media for describing surrounding world. Fine-grained visual explanations are helpful for people to understand the reasons or motivation of vision system when it makes classification decision. Base on the pioneer work of Lisa, this paper proposes a new model for discriminative visual explanation generation. It extracts res5c image features from deep residual network and uses multimodal compact bilinear strategy for multimodal information fusion. Selective attention mechanism is introduced to focus on visual parts that are most related to the predicted category information. The proposed network both considers spatial distribution of image content and fusion strategy that better model different modal information. The result on CUB Bird Dataset shows that our model can improve the quality of the explanation statement, which indicates that our proposed network is effective.

Read the paper · More papers on PaperTik