Recent Development on Extractive Rationale for Model Interpretability: A Survey

Hao Wang, Yong Dou · 2022

Interpretability is an important yet difficult issue for natural language understanding, which aims to understand the behaviour mode of deep neural models, and hence to obtain corresponding explanation for model prediction. For interpretability, extractive rationales are some highlights in the original input, which can sufficiently and comprehensively decide the label, while generative explanations are free texts, which describe the inference process with natural language. In this paper, we focus on extractive rationale, and give a detailed review on recent development. Specifically, firstly we will introduce the task formulation, evaluation metrics and datasets. Then we will introduce several kinds of rationale extraction models, including vanilla attribution algorithm, AA-based training, and select-predict pipeline. And we next list the diverse utilizations of rationale. Finally, we summarize the main challenges of this field.

Read the paper · More papers on PaperTik