Bias Analysis in Language Models using An Association Test and Prompt Engineering
Ravi Varma Kumar Bevara, Ting Xiao, Farahnaz Hosseini, Junhua Ding · 2023
Despite significant progress in the fields of machine learning and deep learning, there remains a sense of mistrust regarding the use of these models in real-world scenarios. This mistrust can be partly attributed to semantic biases in text, especially within the realm of commercial natural language processing (NLP). In this work, we analyze genre bias in movie reviews using the Word Embedding Association Test (WEAT). We compare bias across foundational transformer models, including BERT, DistilBert, RoBERTa, T5, XLNet, and GPT2, along with traditional approaches like Glove and Word2Vec. Our analysis shows that while the underlying data contains bias, different models exhibit varied bias levels due to their distinct architectures and training objectives. To mitigate bias, we propose a simple yet effective prompt engineering technique. Incorporating prompts led to a noticeable reduction in bias across different genres, with the effect sizes indicating that using prompts decreased bias by approximately 35% on average compared to scenarios without prompts. Our work provides new analysis that sheds light on prompt engineering techniques to address the pressing issue of semantic bias in NLP models. We believe continued research in this direction can lead to more transparent and fair AI systems.