Fully Auto-Regressive Multi-modal Large Language Model for Contextual Emotion Recognition

Hiroshi Nonaka, Damian Valles · 2024

Large Language Models (LLMs) have gained popularity due to their high performance in natural language processing. This capability is underpinned by their ability to contextualize relations within text data. The application of LLMs also extends to multi-modal tasks, such as image captioning and video analysis. Research endeavors have been undertaken to leverage LLMs for emotion recognition tasks by utilizing visual, audio, and text data. Inspired by natural human approaches to emotion recognition, we propose a simple yet powerful, multi-modal LLM architecture for emotion recognition in conversations (ERC). Our proposed framework aims to contextualize the change of emotions better. We developed the Fully Auto-Regressive multi-modal LLM for Contextual Emotion Recognition (FARCER) based on this idea. This model consists of the instruction-tuned LLaMA3, a vision encoder, and linear layers that map visual embeddings to the LLM input space. FARCER significantly improved ERC benchmarks over the uni-modal LLaMA3-Instruct, although the LLM and vision encoder’s parameters remained frozen during training. Fine-tuned FARCER demonstrated high performance comparable to other state-of-the-art (SOTA) models, highlighting the potential of our context-focused design combined with conversational LLMs for ERC.

Read the paper · More papers on PaperTik