BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
Zhiting Fan, Ruizhe Chen, Ruiling Xu, Zuozhu Liu · 2024
Evaluating the bias in Large Language Models (LLMs) becomes increasingly crucial with their rapid development.However, existing evaluation methods rely on fixed-form outputs and cannot adapt to the flexible open-text generation scenarios of LLMs (e.g., sentence completion and question answering).To address this, we introduce BiasAlert, a plug-and-play tool designed to detect social bias in open-text generations of LLMs.BiasAlert integrates external human knowledge with inherent reasoning capabilities to detect bias reliably.Extensive experiments demonstrate that BiasAlert significantly outperforms existing state-of-theart methods like GPT4-as-A-Judge in detecting bias.Furthermore, through application studies, we demonstrate the utility of BiasAlert in reliable LLM bias evaluation and bias mitigation across various scenarios.Model and code will be publicly released.