Human-AI Collaboration Enables GPT-4 to Achieve Human-Level User Feedback in Emotional Support Conversations: Integrative Modeling and Prompt Engineering Approaches (Preprint)
Yinghui Huang, Lie Li, Wanghao Dong, Yuhang Dong, Yingdan Huang, Hui Liu · 2024
BACKGROUND Emotional support is crucial in enhancing social interactions, facilitating psychological interventions, and improving customer service outcomes by addressing individuals' emotional needs. The emergence of large language models (LLMs) offers potential for delivering emotional support on a large scale, but their effectiveness compared to human counselors has not been well understood. Evaluating and enhancing the emotional support capabilities of LLMs through targeted user-centered strategies is crucial for their successful real-world integration. OBJECTIVE This study aims to evaluate the emotional support capabilities of LLMs, specifically GPT-4o, and to introduce an integrative automatic evaluation framework focused on user perceived feedback (UPF). The framework seeks to enhance LLM performance in emotional support conversations (ESCs) by identifying psycholinguistic clues as intrinsic evaluation metrics and utilizing a customized Chain-of-Thought (CoT) prompting strategy. METHODS The study utilized a dataset of ESCs from human counselors to develop an explanatory predictive model using explainable artificial intelligence methods, following an integrative modeling paradigm rooted in computational social science. This model was designed to evaluate and interpret UPF scores for GPT-4o. Additionally, Hill's three-stage model of helping was integrated into a manually customized CoT prompting framework to evaluate GPT-4o's performance in ESCs. RESULTS GPT-4o achieved high UPF scores, demonstrating relative stability in performance, but it still significantly lags behind human counselors overall (Cliff's Delta = 0.087, P < 0.001). The evaluation framework identified 41 distinct linguistic clues related to emotional expression, social dynamics, cognitive processes, linguistic style, and decision-making stages, enhancing the understanding of both processes and outcomes in ESCs. Notably, GPT-4o's UPF scores significantly improved with the use of manually customized COT prompts (Cohen's d = 0.378, P < 0.001), showing no significant difference from the average performance of human counselors overall (Cliff's Delta = -0.014, P= 0.47). However, the COT prompts demonstrated a considerable advantage in specific emotion categories such as fear (Cliff's Delta = -0.23, P = 0.002), sadness (Cliff's Delta = -0.105, P = 0.012), and issues related to breakups with partners (Cliff's Delta = -0.06, P = 0.254). Compared to human counselors, GPT-4o is effective in reducing negative language and conveying emotional tone, but its overemphasis on emotional content weakens its causal reasoning, engagement prompting, and cognitive depth, limiting its ability to handle complex questions and scenarios. CONCLUSIONS This study offers preliminary evidence of GPT-4o's emotional support capabilities and introduces a UPF-centered integrative evaluation framework for ESCs. The findings suggest a cautiously optimistic outlook for applying advanced LLMs in emotional support services, though significant challenges persist, particularly in deepening conversational exploration and personalizing language. The proposed framework emphasizes the integration of human expertise into LLMs, enhancing their efficacy and contributing to developing trustworthy AI-based emotional support services.