Guardrails, not Guidance: Understanding Responses to LGBTQ+ Language in Large Language Models

Joshua Tint · 2025

Language models have integrated themselves into many aspects of digital life, shaping everything from social media to translation.This paper investigates how large language models (LLMs) respond to LGBTQ+ slang and heteronormative language.Through two experiments, the study assesses the emotional content and the impact of queer slang on responses from models including GPT-3.5, GPT-4o, Llama2, Llama3, Gemma and Mistral.The findings reveal that heteronormative prompts can trigger safety mechanisms, leading to neutral or corrective responses, while LGBTQ+ slang elicits more negative emotions.These insights punctuate the need to provide equitable outcomes for minority slangs and argots, in addition to eliminating explicit bigotry from language models.

Read the paper · More papers on PaperTik