A Hypergraph Based Contextual Relationship Modeling Method for Multimodal Emotion Recognition in Conversation

Nannan Lu, Zhiyuan Han, Zhen Tan · IEEE Transactions on Multimedia · 2024

Emotion Recognition in Conversation (ERC) has gained considerable attention due to its importance in human-computer interaction. In ERC task, the combination of multimodal information and contextual information is necessary since it can help the model understand emotional changes in the context from multiple perspectives. As Graph Neural Networks (GNNs) have shown the superiority in relation modeling, many graph-based methods have been proposed to improve the performance of emotion recognition by utilizing the edges to mine the contextual relationship and multimodal relationship in a conversation. However, the existence of numerous redundant edges and excessively complex modality interaction in the graph hinders the model from capturing the truly effective dependency information for emotion recognition. In this paper, we propose a Hypergraph based Contextual Relationship Modeling Method (HyperCRM) to carry out the ERC task. HyperCRM models a conversation as a hypergraph instead of a graph, which defines two types of hyperedges, namely speaker-level hyperedge and sequence-level hyperedge, to represent the contextual relationship within the same speaker and the local sequence of the conversation, respectively. Multimodal information is leveraged here as the node feature representation by the feature concatenation. In addition, an improved hypergraph convolution method is designed to capture the long-range contextual information by three-stage information propagation in the hypergraph, including node-hyperedge, hyperedge-hyperedge and node-hyperedge. The extensive experiments on two public datasets shows the new State-Of-The-Art (SOTA) results, to further demonstrate that the proposed method can simply make use of the multimodal information and effectively model the complex contextual relationships in the conversation.

Read the paper · More papers on PaperTik