A Comprehensive Review of Bias Detection and XAI Techniques in NLP
Loganathan Bhavaneetharan · 2025
Bias in Natural Language Processing (NLP) models pose urgent ethical and societal challenges by perpetuating stereotypes and inequalities. This review provides a comprehensive overview of state-of-the-art Explainable AI (XAI) techniques—such as LIME, SHAP, and Integrated Gradients—for detecting and mitigating bias. We detail our methodology for selecting relevant studies and analyze key NLP datasets (e.g., StereoSet, CrowS-Pairs) to uncover specific limitations like language coverage and intersectional gaps. Moreover, we highlight emerging challenges in scalability, cultural diversity, and regulatory compliance. By integrating these findings, we propose actionable strategies for fair and transparent NLP systems, positioning this work as a foundation for ongoing improvements in equitable AI solutions.