Auditing Black-Box AI Systems Using Counterfactual Explanations
Murali Krishna Pasupuleti · International Journal of Academic and Industrial Research Innovations(IJAIRI) · 2025
Abstract: The widespread deployment of black-box artificial intelligence (AI) systems in high-stakes domains such as healthcare, finance, and criminal justice has intensified the demand for transparent and accountable decision-making. This study investigates the utility of counterfactual explanations as a method for auditing opaque AI models. By answering the question of what minimal change in input would alter a model’s output, counterfactuals offer a practical and interpretable means of understanding algorithmic decisions. The methodology employs Random Forest and Neural Network classifiers on two benchmark datasets—German Credit (finance) and MIMIC-III (healthcare)—to examine prediction outcomes before and after applying counterfactual explanations. Techniques including DiCE (Diverse Counterfactual Explanations) and gradient-based methods are evaluated using fidelity, sparsity, proximity, and fairness metrics. Results indicate that counterfactual interventions significantly reduce statistical disparity while maintaining high predictive accuracy, with interpretability scores improving by an average of 22%. Additionally, retraining models using counterfactual insights leads to enhanced fairness without notable performance degradation. The findings underscore the potential of counterfactual auditing to uncover bias, enhance transparency, and facilitate compliance with emerging AI governance standards. This research contributes to the development of interpretable, ethically aligned AI systems suitable for critical decision-support environments. Keywords: Counterfactual explanations, black-box AI, model auditing, interpretability, fairness, algorithmic transparency, DiCE, AI governance, ethical AI, decision-support systems