Embodied Conversational Agents: Deep Learning Based Multimodal Sentiment Analysis
Andrei Valerievich Nosov, Sergey Evgenievich Shtekhin, Alexander A. Lyubchenko · 2024
This paper aims to delve into the realm of Embodied Conversational Agents (ECAs) with integrated deep learning and propose a novel concept of multimodal Embodied Conversational Agent. Our approach extends beyond traditional conversational agents (CAs), encompassing object-centric and human-centric reasoning, influenced by the MBTI methodology. We propose a novel two-stage approach. The first stage involves predicting the MBTI personality through multimodal analysis, utilizing these predictions as proxies for personality annotations to enhance reasoning capabilities. "Reasoning capabilities" refers to an expanded understanding of the concepts of perceiving and using from the MSCEIT methodology. This specifically includes the recognition of facial expressions, speech, and text, as well as the classification of traits within the MBTI framework. Additionally, it involves adjusting the conversational agent's characteristics to match those recognized in human reactions. By integrating MBTI personality model frameworks with NLP and CV approaches, we aim to offer an original solution for advancing multimodality in the context of ECAs. Finally, we seek to present conceptual evidence and further real-world applications to support the viability of our approach, solidifying its potential impact on the field of multimodal sentiment analysis and conversational agent technology. The research outcome is the development of a comprehensive model for effective communication with conversational AI agents.