RESPECT: A framework for promoting inclusive and respectful conversations in online communications
Shaina Raza, Abdullah Yahya Mohammed Muaad, Emrul Hasan, Muskan Garg, Zainab Al-Zanbouri, Syed Raza Bashir · Natural Language Processing Journal · 2025
Toxicity and bias in online conversations hinder respectful interactions, leading to issues such as harassment and discrimination. While advancements in natural language processing (NLP) have improved the detection and mitigation of toxicity on digital platforms , the evolving nature of social media conversations demands continuous innovation. Previous efforts have made strides in identifying and reducing toxicity; however, a unified and adaptable framework for managing toxic content across diverse online discourse remains essential. This paper introduces a comprehensive framework R ESPECT designed to effectively identify and mitigate toxicity in online conversations. The framework comprises two components: an encoder-only model for detecting toxicity and a decoder-only model for generating debiased versions of the text. By leveraging the capabilities of transformer-based models, toxicity is addressed as a binary classification problem. Subsequently, open-source and proprietary large language models are utilized through prompt-based approaches to rewrite toxic text into non-toxic, and making sure these are contextually accurate alternatives. Empirical results demonstrate that this approach significantly reduces toxicity across various conversational styles, fostering safer and more respectful communication in online environments.