Can Causal DAGs Generate Data-based Explanations of Black-box Models?
Arman Ashkari, El Kindi Rezig · 2024
AI models (e.g., machine learning or deep learning models) have become ubiquitous in our everyday life. This led to an ever-increasing complexity of these systems to accommodate various use-cases and domains (e.g., ML for healthcare). As a result, explaining the output of these systems has become crucial in many applications such as ML fairness, and data debugging. Existing explainable AI (XAI) techniques can be categorized into two categories: (1) Model-agnostic approaches (e.g., LIME, Shapley values) that quantify the feature importance for a given prediction. (2) White-box XAI techniques which assume knowledge of the model’s function and parameters (e.g., influence functions) and can quantify the impact of removing individual training points on the model’s parameters without re-training. With the rise in complexity of ML models, tracing their predictions to training data without re-training has become ineffective. As a result, model-agnostic XAI methods have become the de facto methods to explain black-box models. However, there is no model-agnostic approach that can link a prediction made by the model to subsets of the training data that are necessary to produce it. We present CausalExplain, a work-in-progress system that combines adversarial training and causal reasoning to produce the top-k training data subsets that are most responsible for a given prediction made by a black-box model. CausalExplain only interacts with the underlying model through its input-output interface and assumes no knowledge of the model’s function or parameters.