Multi-Dimensional Gender Bias Classification
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, Adina Williams · 2020
Machine learning models are trained to find patterns in data.NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.In this work, we propose a novel, general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.In addition, we collect a new, crowdsourced evaluation benchmark.Distinguishing between gender bias along multiple dimensions enables us to train better and more fine-grained gender bias classifiers.We show our classifiers are valuable for a variety of applications, like controlling for gender bias in generative models, detecting gender bias in arbitrary text, and classifying text as offensive based on its genderedness.