Debiasing Gender biased Hindi Words with Word-embedding
Arun K. Pujari, A.P. Mittal, Anshuman Padhi, Anshul Jain, Mukesh Kumar Jadon, Vikas Kumar · 2019
Word-embedding is a major machine learning technique for computational applications of languages. For a given corpus, the process of word-embedding is to embed each word onto multi-dimensional space such that semantic similarities between similar words are retained. While learning the similarity as encapsulated in the training corpus, the embedding process inadvertently captures many other inherent features present in the corpus. One such thing is the bias arising out of stereotyping present in almost all the corpus no matter how extensively used and trusted they are. We study this aspect of word-embedding in the context of Hindi language. We show that many gender-neutral words in Hindi are mapped to vectors which are inclined towards one gender or the other in multi-dimensional space. We propose a new algorithm of debiasing and demonstrate its efficacy in the context of Hindi language. Further, we build a SVM-based classifier that determines whether a gender-neutral word is classified as neutral or otherwise. We corroborate our claim with experimental results on large number of individual words. This work is first ever result on debiasing in Hindi Language and our new debiasing algorithm can be applicable in the context of any language.