Response Generation through Social Reasoning in Large Language Models with Direct Diverse Preferences Optimization
Maryam Amirizaniani, Elias Martin, Afra J. Mashhadi, Chirag Suresh Shah · 2025
Large Language Models (LLMs) have demonstrated capabilities across a wide range of information retrieval (IR) tasks, including generating reasoning-based responses for social questions. A common approach to enhancing these abilities involves optimizing model behavior based on human preferences through an implicit reward function. However, existing frameworks are often limited to binary win-lose output optimization, which fails to capture the nuanced and diverse nature of human preferences, especially in socially complex scenarios that involve trade-offs and multiple valid viewpoints. To address this limitation, we propose Direct Diverse Preference Optimization (DDPO), a novel framework that models user preference behavior by leveraging ranked sets of both preferred (win) and non-preferred (loss) responses for social reasoning task. DDPO aims to increase the probability of higher-ranked win responses while decreasing the probability of lower-ranked loss responses, effectively capturing the graded nature of human preferences. This rank-aware optimization enables LLMs to generate reasoning responses that are more closely aligned with varied human preferences. We evaluate DDPO on real-world social questions collected from Reddit and Lemmy using two LLMs. The results show an average improvement of 8.7% in reasoning-based logical inference metrics, demonstrating its effectiveness in generating responses that better reflect human expectations. The novelty of this work lies in leveraging user modeling through ranked and diverse human preferences, derived from real-world social platforms, to guide the generation of socially grounded reasoning responses.