Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, Helen M. L. Meng · 2022
Warning: this paper contains content that may be offensive or upsetting.Among the safety concerns that hinder the deployment of open-domain dialog systems (e.g., offensive languages, biases, and toxic behaviors), social bias presents an insidious challenge.Addressing this challenge requires rigorous analyses and normative reasoning.In this paper, we focus our investigation on social bias measurement to facilitate the development of unbiased dialog systems.We first propose a novel DIAL-BIAS FRAMEWORK for analyzing the social bias in conversations using a holistic method beyond bias lexicons or dichotomous annotations.Leveraging the proposed framework, we further introduce the CDIAL-BIAS DATASET which is, to the best of our knowledge, the first annotated Chinese social bias dialog dataset.We also establish a finegrained dialog bias measurement benchmark, and conduct in-depth analyses to shed light on the utility of detailed annotations in the proposed dataset.Lastly, we evaluate several representative Chinese generative models using our classifiers to unveil the presence of social bias in these systems. 1