Uncovering the Linguistic Roots of Bias: Insights and Mitigation in Large Language Models
Lauren Benson, Ahmet Okutan, Roopa Vasan · 2025
The complexity, diversity, and opacity of large language models (LLMs) present significant challenges in ensuring fairness and trustworthiness.While existing research has made strides in creating interpretable and controllable models, there remains a critical need for generalizable and real-world-applicable bias mitigation solutions.In this work, we propose a novel methodology that bridges the gap between linguistic capabilities and biased behavior, uncovering and leveraging hidden correlations to mitigate bias holistically.By analyzing the interplay between unrelated linguistic features and bias benchmarks, we identify natural language understanding (NLU) tasks that significantly influence biased behavior.This insight enables the creation of bias-mitigated models through a multi-task learning framework that augments behavior with minimal performance trade-offs.Our findings reveal that (1) specific NLU capabilities correlate with multiple perspectives of biased behavior, and (2) these relationships can be applied to mitigate bias without sacrificing model accuracy.In a real-world domain-specific application, according to the Aequitas disparity measures using client-provided demographic labels, our method reduced gender and age bias by 14% and 11%, respectively, with only a 0.03% decrease in Micro-F1 score.This approach introduces a scalable, generalized pathway for addressing systemic biases in LLMs, demonstrating the potential to improve fairness and trustworthiness in AI systems across diverse and sensitive applications.