A Two-Stage Approach for Implicit Bias Detection in Generative Language Models

Jeremy Edwards, Renjie Hu, Amaury Lendasse, Alexander Schlager, Peggy Lindner · 2024

Machine learning and AI are increasingly popular for their impressive task performance.Yet, Natural Language Processing (NLP) models often inadvertently learn harmful biases related to gender and race, leading to skewed predictions.Literature distinguishes between direct and indirect bias.Current research aims to detect and mitigate these biases in machine learning models.This study introduces a two-stage approach to identify both types of gender bias in generative large language models (LLMs), confirming that they can manifest both direct and indirect biases.

Read the paper · More papers on PaperTik