ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Lin Zi, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, Jingbo Shang · 2023

Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays.However, previous efforts in toxicity detection have been mostly based on benchmarks derived from social media content, leaving the unique challenges inherent to real-world user-AI interactions insufficiently explored.In this work, we introduce TOXICCHAT, a novel benchmark based on real user queries from an open-source chatbot.This benchmark contains the rich, nuanced phenomena that can be tricky for current toxicity detection models to identify, revealing a significant domain difference compared to social media content.Our systematic evaluation of models trained on existing toxicity datasets has shown their shortcomings when applied to this unique domain of TOXI-CCHAT.Our work illuminates the potentially overlooked challenges of toxicity detection in real-world user-AI conversations.In the future, TOXICCHAT can be a valuable resource to drive further advancements toward building a safe and healthy environment for user-AI interactions.

Read the paper · More papers on PaperTik