Investigating Automated Mechanisms for Multi-Modal Prediction of User Online-Video Commenting Behaviour

Hao Wu, François Pitié, Gareth J. F. Jones · 2021

Online-video commenting is attracting increasing attention among young people, particularly in the form of "danmu" comments in Asia. These provide a channel for engagement enabling users to share time-synchronous comments on videos with other viewers. Danmu form community discussions of video content and frequently provoke extensive further contributions. The motivation to add danmu comments at specific points in videos are not obvious. In this paper, we explore the potential for predicting user online-video comment distributions using multi-modal signals from the video content stream. To address this task we integrate multiple sources of information, including video frame content, the audio signal, as well as video subtitles, in an end-to-end neural framework. Specifically, text, visual and audio input are encoded respectively and then a transformer framework is used to learn and combine attention aware representation of three modalities. We evaluate the system using retrieval-based evaluation metrics, including mean average precision (mAP) and normalized discounted cumulative gain (NDCG). We conduct experiments on an expanded publicly available danmu commenting dataset. Our model significantly outperformed an LSTM multi-modal baseline method.

Read the paper · More papers on PaperTik