Recommendation of Hashtags Using Deep Learning Methods Based on Multimodal Data
Sergiy Yakovlev, Nataliia Shapoval · Artificial Intelligence · 2024
Generating image text captions is an important task and aims to automatically generate a text description for an image. Recommendation of hashtags is a practical option for this task. Hashtags contribute to increasing the relevance of content for the audience and ensure better visibility of publications. The problem of choosing optimal hashtags becomes especially relevant for social platforms, where users generate huge amounts of content with different types of modalities — images, text captions, videos, etc. There are a number of challenges that need to be addressed when solving this problem. First, text captions for posts are often short or even absent. Secondly, multimodal algorithms often do not take into account the previous activity of the user, which can significantly limit the quality of recommendations. Third, the balance between the importance of textual and visual cues may vary depending on the nature of the publication. The purpose of this study is to develop a modified feature fusion algorithm for the task of multimodal hashtag recommendation, which is able to take into account the context of the user's previous history, adaptively evaluate the importance of textual and visual features, and improve the quality of recommendations in cases of the absence or weakness of textual description. As part of the study, a model was modified that, in addition to analyzing the image and text caption, takes into account part of the previous history of user interactions. The main contribution is a new feature fusion module that weights their importance depending on the context. This approach allows to improve the relevance of recommendations in situations where the textual modality is not informative enough, which is a common problem in real data. The experimental results confirmed that the proposed feature fusion module provided more accurate hashtag recommendations, especially for cases with short or missing text captions