You Type a Few Words and We Do the Rest: Image Recommendation for Social Multimedia Posts
Tianlang Chen, Yuxiao Chen, Han Guo, Jiebo Luo · 2018
In this paper, we introduce a new application that can be employed on many social media platforms. We intend to recommend related images from local (e.g. user's local mobile phone storage) and global (e.g. platform's server) image pools while a user is composing the text to post a status. To make the recommendation system applicable to different platforms with or without the support of online computing, we propose two independent frameworks that recommend images at image level and data-driven category level based on a text, respectively. For image-level recommendation, our framework recommends images for a text by predicting the affinity scores of image-text pairs and recommending the images with the highest scores for the text. To improve the ranking performance, we propose a novel patch-level image-text matching framework which strikes a balance between local and global matching of image-text pairs. In particular, it first extracts the affinity between each local word and image patch, then leverages different kinds of attention mechanisms to respectively weight the local words and patches for computing the final image-text affinity scores. For category-level recommendation, we first classify images into categories in an unsupervised way, and then propose a multi-task LSTM-based framework with effective user feature for recommendation. Extensive experiments in two real-world social media datasets demonstrate the effectiveness of the proposed models, which significantly outperform the baselines. We also visualize the patch-word matching details to provide an insight into the image-level recommendation framework and demonstrate the strong capacity of category-level recommendation framework to recommend images in an online fashion.