Personalized Visual Emotion Classification via In-context Learning in Multimodal LLM

Ryo Takahashi, Naoki Saito, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama · 2024

This paper presents a personalized method for visual emotion classification by in-context learning in a Multimodal Large Language Model (MLLM). Although MLLM has accumulated a vibrant store of knowledge and is expected to be used for visual emotion classification tasks, a problem exists that makes it difficult to accurately classify different emotions for different users for the same visual stimulus. To deal with this problem, the proposed method performs the classification of different emotions for each user by in-context learning in MLLM. By providing each user with a prompt that summarizes the emotion elicited by the visual stimulus and the reason for the emotion, the proposed method has made it possible to classify emotions using MLLM, which considers each user’s unique information. Experimental results show the effectiveness of the proposed method, which introduces in-context learning in MLLM to the visual emotion classification.

Read the paper · More papers on PaperTik