Multi-modal guided attention for live video comments generation
Yuchen Ren, Yuan Yuan, Lei Chen · International Conference on Computer Graphics, Artificial Intelligence, and Data Processing (ICCAID 2021) · 2022
With the blooming of online video applications, live commenting is an emerging feature of online video sites. The live video comments generation (LVCG) task aims to generate live comments for videos while considering both the video and the surrounding comments made by other viewers. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To overcome the problem of insufficient multimodal interactions for live video comments generation, we built two basic attention blocks: the self attention (SA) block that can model the dense intramodal interactions; and the x-guided attention (XGA) block to model the dense intermodal interactions. After that, by modular compositions of the SA and XGA blocks, we propose different multimodal transformer architectures to handle the multimodal features. Finally, experiments show that our proposed multimodal guided attention models significantly outperform previous methods in most of the metrics.