Performance Evaluation Metrics for Empathetic LLMs
Yuna Hong, Bonhwa Ku, Hanseok Ko · Information · 2025
With the rapid advancement of large language models (LLMs), recent systems have demonstrated increasing capability in understanding and expressing human emotions. However, no objective and standardized metric currently exists to evaluate how empathetic an LLM’s response is. To address this gap, we propose a novel evaluation framework that measures both sentiment-level and emotion-level alignment between a user query and a model-generated response. The proposed metric consists of two components. The sentiment component evaluates overall affective polarity through Sentlink and the naturalness of emotional expression via NEmpathySort. The emotion component measures fine-grained emotional correspondence using Emosight. Additionally, a semantic component, based on RAGAS, assesses the contextual relevance and coherence of the response. Experimental results demonstrate that our metric effectively captures both the intensity and nuance of empathy in LLM-generated responses, providing a solid foundation for the development of emotionally intelligent conversational AI.