Detecting Bias in LLMs' Natural Language Inference Using Metamorphic Testing
Zhehao Li, Jinfu Chen, Haibo Chen, Ling Xu, Wuhao Guo · 2024
The increasing integration of Large Language Mod-els (LLMs) into pivotal decision-making contexts highlights the necessity for rigorous scrutiny to ensure fairness, especially in tasks involving Natural Language Processing (NLP). This paper introduces a novel method employing Metamorphic Testing (MT) to discern bias within the Natural Language Inference (NLI) task performed by LLMs. Our approach establishes a structured framework for generating test cases using Metamorphic Relations (MRs), incorporating nuanced demographic attributes such as sex, race, occupation, age, and socioeconomic status into textual inputs. We conduct prelimi-nary experiment with 3 common datasets, over 400k test cases in 4 state-of-art LLMs. Through evaluation of model predictions across these MRs, we unveil potential biases indicative of fairness violations. Preliminary investigations conducted with prominent LLMs underscore the effectiveness of our method in illuminating disparities in model performance, thus providing valuable insights into the bias implications of LLMs in NLI tasks. Our method is intuitive and adaptable for other NLP tasks. To the best of our knowledge, this represents the first application of MT to NLI bias detection.