Adversarial Debiasing using Regularized Direct Entropy Minimization in Wasserstein GAN with Gradient Penalty
Md Nur Amin, Abdullah Al Imran, Fatih S. Bayram, Alexander Jesser · 2023
The promise of machine learning algorithms are overshadowed by the ethical challenges they present, especially when biases in data lead to discriminatory outcomes based on gender or other sensitive attributes. This paper introduces an innovative methodology to address gender bias in machine learning models, particularly when processing textual data. By harnessing the power of adversarial training with the Wasserstein GAN enhanced by gradient penalty (WGAN-GP), our approach strives to minimize gender-specific information in the model’s latent representations. Central to our methodology is the integration of a Direct Entropy Minimization (DEM) regularizer, which further ensures that these representations are devoid of gender biases. The dual objective, achieved through our primary classifier and its adversarial counterpart, not only maintains high classification accuracy but also substantially eliminates gender bias.