A residual merged neutral network for multimodal sentiment analysis

Nan Xu, Wenji Mao · 2017

With the continuous development of social networking sites, the volume of social media data has exploded and the user-generated content is becoming more and more diverse. As a result, the modality of massive social media data is no longer confined to the single text mode. This brings new challenges to social media analytics in general and its examplar field such as sentiment analysis in particular. Multimodal sentiment analysis has become an increasingly important research topic in recent years, especially in the context of social media big data. Most of the previous work only focuses on single modality content such as text, image or speech. Moreover, as the traditional sentiment analysis methods often lack the support of scalable deep models, this hinders their usage in processing large amount of online data. To overcome the limitations in the previous work, in this paper, we propose an end-to-end framework for multimodal sentiment analysis based on deep neural network. We propose a Merged Neural Network (MNN) model that utilizes CNNs to extract representations of text and image respectively. To fuse the multimodal features, we introduce the residual model and propose two combined merged strategies, namely the Early-RMNN (i.e. Early Residual MNN) and Late-RMNN (i.e. Late Residual MNN), to get deeper and more discriminative features than the previous methods. The experiments on two public available datasets demonstrate the effectiveness of our models for multimodal sentiment analysis in comparison with the related methods.

Read the paper · More papers on PaperTik