Inferring Sensitive Attributes from Model Explanations
Vasisht Duddu, Antoine Boutet · Proceedings of the 31st ACM International Conference on Information & Knowledge Management · 2022
Model explanations provide transparency into a trained machine learning model's blackbox behavior to a model builder. They indicate the influence of different input attributes to its corresponding model prediction. The dependency of explanations on input raises privacy concerns for sensitive user data. However, current literature has limited discussion on privacy risks of model explanations.