Individual CNN Hidden-Layer Neurons Are Good Concept Encoders
Abhilekha Dalal, Rushrukh Rayan, Adrita Barua, Samatha Ereshi Akkamahadevi, Md Kamruzzaman Sarker, Cara Widmer, Pascal Hitzler, Eugene Y. Vasserman · Frontiers in artificial intelligence and applications · 2025
This work discusses some recent advances and continued limitations in explainable deep learning, leveraging neurosymbolic methods to identify human-understandable concepts which are “recognized” by machine models and drive their behaviors. Results suggest that models’ hidden layers do encode information as discrete human-understandable concepts, and that the information is meaningful not just from the point of view of the model output, or even groups of hidden-layer neurons, but even down to individual neurons.