Leading and non-leading prompts: Quantifying gender bias in Large Language Models through BiasBloom corpus
Victoria Muñoz-García, Juan Pablo Consuegra‐Ayala, Paloma Moreda · Knowledge-Based Systems · 2025
The increasing use of Artificial Intelligence by non-experts underscores the need to address biases in its outputs, particularly in gendered languages like Spanish, where linguistic features reflect gender. This paper proposes a validated methodology to quantify gender bias in text generated by Large Language Models (LLMs) in Spanish. The approach involves creating gendered seed-word lists, building a Spanish-specific corpus using curated prompts, and analyzing gender polarity and co-occurrences within the generated text. Validated using five state-of-the-art LLMs—GPT-3.5, GPT-4o, Llama 3, Gemini 1.5, and Mixtral8x7b—this study provides a systematic framework for bias detection in Spanish and highlights differences in model performance. By addressing the challenges of bias in gendered languages, this research aims to contribute to the development of equitable AI systems and advances methodologies for bias quantification. The findings highlight gender disparities in LLMs, with masculine and feminine biases shifting depending on the discourse structure. Specifically, the analysis reveals a prevalence of masculine biases in non-leading prompts and feminine biases in leading prompts, underscoring the nuanced ways in which prompt positioning can influence gender representation in model outputs.