Identifying and Mitigating Gender Bias in Language Models: A Fair Machine Learning Approach
Sangeeth Ajith, M Rithani, R S SyamDev · 2023
Large language models (LLMs) in natural language processing (NLP) have shown remarkable performance but are marred by biases inherited from the data they are trained on. Gender bias, in particular, poses a significant challenge, perpetuating societal inequalities. Existing debiasing methods often compromise performance and lack generalizability. To address this, we propose an adversarial debiasing method for transformer-based models, demonstrated on the BERT architecture. Our approach showcases the effectiveness and computational efficiency in debiasing pre-trained transformers for autocompletion tasks. We evaluate extrinsic fairness measures and demonstrate improved sentiment fairness, maintaining model performance as indicated by perplexity measurements. Additionally, we illustrate the transferability of fairness improvements to downstream tasks without compromising performance. The proposed method leads to an increase in accuracy, with the debiased model showing a 3% improvement in accuracy on the SemEval dataset and a 4% improvement on the Reddit dataset. Our findings emphasize the critical need to prioritize gender bias mitigation for more ethical and inclusive language processing, promoting equitable AI systems.