AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark
Abhay Gupta, Ece Yurtseven, Philip Meng, Kevin Zhu · 2024
Detecting biases in natural language understanding (NLU) for African American Vernacular English (AAVE) is crucial to developing inclusive natural language processing (NLP) systems.To address dialectinduced performance discrepancies, we introduce AAVENUE (AAVE Natural Language Understanding Evaluation), a benchmark for evaluating large language model (LLM) performance on NLU tasks in AAVE and Standard American English (SAE).AAVENUE builds upon and extends existing benchmarks like VALUE, replacing deterministic syntactic and morphological transformations with a more flexible methodology leveraging LLMbased translation with few-shot prompting, improving performance across several evaluation metrics when translating key tasks from the GLUE and SuperGLUE benchmarks.We compare AAVENUE and VALUE translations using five popular LLMs and a comprehensive set of metrics including fluency, BARTScore, quality, coherence, and understandability.Additionally, the fluency of AAVENUE is validated by annotations from AAVE speakers.Our evaluations reveal that LLMs consistently perform better on SAE tasks than AAVEtranslated versions, underscoring inherent biases and highlighting the need for more inclusive NLP models.We have open-sourced our source code on GitHub and created a website to showcase our work at https://aavenue.live.