Explaining What a Neural Network has Learned: Toward Transparent Classification
Kasun Amarasinghe, Milos Manic · 2019
Deep Neural Networks (DNNs) have limited ability to explain their acquired knowledge or decision rationale. As a result, end-users perceive DNNs as black-boxes and are hesitant to fully adopt them in safety-critical applications. Therefore, developing explainable DNNs has become a prime interest in neural network research. This paper presents a methodology for linguistically explaining the knowledge a DNN classifier has acquired in training. The main objective is to help users understand what the DNN has learned about each class. The presented methodology is fuzzy logic based and involves end-users of the system in the explanation process, enabling users to customize the explanations to match their requirements. This paper presents the explanation methodology, metrics of explanation quality, validation steps, and a discussion of advantages and limitations. The explanation methodology was implemented on a benchmark classification problem. Experimental results demonstrated the method's capability to explain the DNN-knowledge and validated the explanations.