Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

Youngjun Lee, Dokyong Lee, Junyoung Youn, Kyeong-Jin Oh, Byungsoo Ko, Jonghwan Hyeon, Ho‐Jin Choi · 2024

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools.However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term social interaction, and (2) a lack of personalized image-sharing behavior.In this work, we introduce STARK , a large-scale long-term multi-modal conversation dataset that covers a wide range of social personas in a multi-modality format, time intervals, and images.To construct STARK automatically, we propose a novel multi-modal contextualization framework, MCU, that generates longterm multi-modal dialogue distilled from Chat-GPT and our proposed Plan-and-Execute image aligner.Using our STARK, we train a multimodal conversation model, ULTRON 7B, which demonstrates impressive visual imagination ability.Furthermore, we demonstrate the effectiveness of our dataset in human evaluation.We make our source code and dataset publicly available 1 .

Read the paper · More papers on PaperTik