Mitigating Bias in Large Language Models Leveraging Multi-Agent Scenarios
Jens Lünstedt, Tim Schlippe · 2025
Mitigating bias in large language models (LLMs) is essential, as biased outputs can perpetuate harmful stereotypes and negatively influence decision-making [1]. LLM-based multiagent scenarios, which have gained attention for their ability to simulate human-like collaboration in tasks such as decisionmaking and strategic planning [2], may reduce bias with advisory LLMs, similarly to how human ethics boards maintain fairness. Consequently, we evaluated three LLM-based multi-agent scenarios to mitigate biased responses to human prompts: (1) a single bias expert agent, (2) a team of bias expert agents, and (3) a simulated human ethics board. To measure bias reduction, we used a subset of the BBQ corpus, a bias benchmark corpus for question answering [3], focusing on eight bias types: physical appearance, disabiliy status, age, nationality status, race / ethnicity, sexual orientation, gender identity, and religion. Results show that all three scenarios reduced bias in LLM outputs by over 20% compared to a single-agent approach.