Towards a Better Understanding of Gradient-Based Explanatory Methods in NLP
Qingfeng Du · Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering · 2021
To grasp what makes the deep learning models arrive at a particular prediction, gradient-based explanatory methods have been widely used in Natural Language Processing (NLP) recently.While the saliency maps of images can be computed directly in the pixel-level input space, the continuous gradient vector for words has to be reduced to a single value to indicate the word-level importance, and existing methods such as Sensitivity Analysis (SA) and Gradient × Input (GI) are either tricky or short of a deep investigation.In this paper, we review the family of gradient-based explanatory methods and discuss their practical implications.Specially, we propose the signed version of GI, namely SignedGI, while some previous work may have misunderstandings on its signedness.We also show the weakness of SA-based methods.We conduct extensive experiments to evaluate these explanatory methods both qualitatively and quantitatively.