Mitigating Gender Bias in Masked Language Models: Introducing Gender Stereotype-Aware Loss Regularizers
Vishal Paul · 2024
The advancement of deep learning-based natural language models has significantly enhanced their capabilities, but it has also amplified the risk of perpetuating inherent biases present in training data. This paper addresses the critical issue of gender bias in Masked Language Models (MLMs), specifically focusing on BERT. I propose two novel loss regularizers, Gender Stereotype-Aware Loss Regularizers (GALs), designed to control and mitigate gender-related stereotypes. The first regularizer (GAL1) minimizes the discrepancy between predictions for original and gender-flipped sentences, while the second (GAL2) ensures balanced prediction probabilities for gendered contexts. By fine-tuning BERT with these regularizers, I demonstrate a reduction in gender bias in its predictions. my methodology includes quantitative and qualitative evaluations using the BUG dataset and other benchmarks to assess the efficacy of my approach. The results indicate that the GAL regularizers lead to a more gender-neutral model, highlighting the potential for improving fairness in NLP systems.