Rather a Nurse than a Physician - Contrastive Explanations under Investigation

Oliver Eberle, Ilias Chalkidis, Laura Cabello Piqueras, Stephanie Brandl · 2023

Contrastive explanations, where one decision is explained in contrast to another, are supposed to be closer to how humans explain a decision than non-contrastive explanations, where the decision is not necessarily referenced to an alternative.This claim has never been empirically validated.We analyze four English text-classification datasets (SST2, DynaSent, BIOS and DBpedia-Animals).We fine-tune and extract explanations from three different models (RoBERTa, GTP-2, and T5), each in three different sizes and apply three post-hoc explainability methods (LRP, GradientxInput, GradNorm).We furthermore collect and release human rationale annotations for a subset of 100 samples from the BIOS dataset for contrastive and non-contrastive settings.A crosscomparison between model-based rationales and human annotations, both in contrastive and non-contrastive settings, yields a high agreement between the two settings for models as well as for humans.Moreover, model-based explanations computed in both settings align equally well with human rationales.Thus, we empirically find that humans do not necessarily explain in a contrastive manner.

Read the paper · More papers on PaperTik