Injecting Bias through Prompts: Analyzing the Influence of Language on LLMs
Nicolás Torres, Catalina Ulloa, Ignacio Araya, Matias Ayala, Sebastián Jara · 2024
Large Language Models (LLMs) have become integral to numerous applications, from natural language processing to automated decision-making. However, these models can exhibit biases that reflect and reinforce societal inequalities. This study investigates how prompts can be strategically designed to inject bias into LLMs, focusing on gender and racial biases. By employing few-shot prompting and varying the verbs and adjectives used in prompts, we explore the extent to which the language of the prompt influences the model's output. Our experiments demonstrate that LLMs can be steered towards generating biased responses by manipulating the prompts they receive. Specifically, we show that prompts containing gender stereotypes or racially charged contexts lead to correspondingly biased model outputs. This research highlights the ethical implications of prompt design in LLMs and underscores the need for robust mechanisms to detect and mitigate bias. Our findings contribute to a deeper understanding of the vulnerabilities of LLMs to prompt-induced biases, providing a foundation for developing fairer and more equitable AI systems.