The model thinks what?! Interpreting deep NLP models with rationales and influence

Sarthak Jain · 2022

They are black-boxes - It is a common criticism leveled at Deep Learning based Predictive Models,especially when deployed in any application where wrong predictions may negatively affect its users. That criticism is correct but is there a solution? More concretely, restricting to supervised learning problems, how can we make a real-valued function with millions of floating-point parameters interpretable; that is, provide an answer to the question "[Why] did a model [M] trained on the example set [D] give input [X] the label [Y]"? What is the definition of Why here? How do we build trust in these models? How can we audit them? Moreover, what recourse do we, as developers or users, have when they invariably return wrong predictions? The goal of this thesis is to answer some of these questions, partially, by formalizing and measuring the effect of the training data [D] and input features [X] on the prediction [Y] for a given model class [M]. Part 1 of this dissertation deals with the issue of incomplete definitions for feature attribution methods. We1 ask whether a commonly used feature attribution method (Attention) reveals the inner workings of a deep neural model, provide a benchmark to test whether these methods faithfully reveal model reasoning, and design an easy-to-use rationalizing model. In Part 2, we focus on improving instance attribution methods by studying faster variants of influence functions and extend these methods to compute token level influence in sequence tagging tasks. Finally, we demonstrate the utility of instance attribution methods for identifying two types of training data issues - artifacts and labeling errors.--Author's abstract

Read the paper · More papers on PaperTik