Neural Network Interpretability: Methods for Understanding and Visualizing Deep Learning Models

A.V.V. Sudhakar, Mansi Sharma, Chinnem Rama Mohan, Er. Suraj Singh, Kuldeep Singh Chouhan, B. T. Geetha · 2024

Neural networks, which are a type of deep learning model, are massively criticized due to their ‘black box’ approach, which does not allow interpreting their decisions. In this research, several post-hoc interpretability techniques, which seek to explain the behavior of such models, are discussed. Specifically, this paper aims to evaluate the application of activation maps, feature importance, and attention mechanisms for explaining CNN and LSTM architectures within two distinct datasets namely Fashion MNIST and IMDB. Thus, the work has a focus on showing that these interpretability methods are not only useful for understanding deep learning models, as well as for reducing bias, fairness, and trust in AI systems. It emphasizes that the complexity and interpretability of models used in socially important tasks should be more thoroughly researched in order to develop effective means of improving artificial intelligence systems while maintaining explanatory power and non-discriminatory approach.

Read the paper · More papers on PaperTik