An Experiment on Localization of Ontology Concepts in Deep Convolutional Neural Networks

Anton Agafonov, Andrew Ponomarev · 2022

Deep neural networks have recently evolved into a powerful AI tool, reaching near-human performance level in many tasks, and in some tasks even surpassing it. However, a significant drawback of neural networks is the lack of explainability and interpretability — it is hard to say why a neural network arrived to a certain conclusion. This significantly limits application of neural networks in critical tasks and undermines trust in human-AI collaboration. It has been recently shown that internal representations constructed by a neural network can often be aligned with a domain ontology. This opens a promising way to provide explanations of a neural network in human terms. In this paper, we discuss the results of the experiment aimed at understanding what layers of a neural network are the most perspective for the alignment with given ontology concept. To do so, we build concept localization maps for XTRAINS — a synthetic dataset consisting of images and their ontological annotations. The importance of such maps is that they can be used for the development of efficient concept alignment heuristics. The experiment mostly supports the intuition that high-level concepts are localized mostly in the activations of last layers of a neural network (near its head), while lower-level concepts might be better extracted from middle layers.

Read the paper · More papers on PaperTik