Employ Multimodal Machine Learning for Content Quality Analysis

Pengfei Du, Xiaoyong Li, Yali Gao · 2020 IEEE 4th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) · 2020

User-generated content varies drastically, and todays mainstream media sites have a lot of information presented in graphic form. So the task of identifying high-quality content becomes increasingly important, and it can improve overall reading time and ctr(click-through rate estimates). Traditional quality assessment method based only on single modes, for example in image quality assessment area it can divide into IQA method and NR-IQA method, and in text quality assessment area we use different indicators such as readability and smoothness, but there are some limitations in analyzing quality from single mode, many information will lose by single modal analysis. In this paper we propose a multimodal quality recognition approach. We use Transformers-XL which is more efficient and accurate for text feature extract, and use inceptionv3 for the original image feature extract, in order to solve the problem of multi-graph fusion we use NeXtVLAD for graph fusion. Then we use siamese network to score the content quality, and the rank loss is used as the optimization objective. We crawled many data from different website such as WeiboTwitter to constitute the validation data set. Compare with other method such as single modal method and other multimodal deep learning method, our approach got a state of art result.

Read the paper · More papers on PaperTik