Learning Fused Representations for Large-Scale Multimodal Classification

Shah Nawaz, Alessandro Calefati, Muhammad Kamran Janjua, Muhammad Umer Anwaar, Ignazio Gallo · IEEE Sensors Letters · 2018

Multimodal strategies combine different input sources into a joint representation that provides enhanced information from the unimodal strategy. In this article, we present a novel multimodal approach that fuses image and encoded text description to obtain an information-enriched image. This approach casts encoded text obtained from Word2Vec word embedding into visual embedding to be concatenated with the image. We employ standard convolutional neural networks to learn representations of information-enriched images. Finally, we compare our approach with the unimodal approach and their combination on three large-scale multimodal datasets. Our findings indicate that the joint representation of encoded text and image in feature space improves the multimodal classification performance aiding the interpretability.

Read the paper · More papers on PaperTik